Search NASA⌕ Search

SEARCH · Search NASA

Results for “Homology modelling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Centrifugal Pump Model for System Codes for Advanced NPP Designs

In this document, a homologous model and one-dimensional line model for centrifugal pump are described to provide the theoretical background of a pump model to be developed in system-level codes. Input and output parameters are proposed, along with a summary of the governing equations for each model. Furthermore, hydraulic performance degradation of centrifugal pump is also modeled in the homologous pump model. This document suggests that system codes provide a homologous and one-dimensional line model for centrifugal pumps, allowing the users to choose the model that best fits their purposes.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

X-ray crystallographic and hydrogen deuterium exchange studies confirm alternate kinetic models for homolog insulin monomers

Despite the crucial role of various insulin analogs in achieving satisfactory glycemic control, a comprehensive understanding of their in-solution dynamic mechanisms still holds the potential to further optimize rapid insulin analogs, thus significantly improving the well-being of individuals with Type 1 Diabetes. Here, we employed hydrogen-deuterium exchange mass spectrometry to decipher the molecular dynamics of newly modified and functional insulin analog. A comparative analysis of H/D dynamics demonstrated that the modified insulin exchanges deuterium atoms faster and more extensively than the intact insulin aspart. Additionally, we present new insights derived from our 2.5 Å resolution X-ray crystal structure of modified hexamer insulin analog at ambient temperature. Furthermore, we obtained a distinctive side-chain conformation of the Asn3 residue on the B chain (AsnB3) by operating a comparative analysis with a previously available cryogenic rapid-acting insulin structure (PDB_ID: 4GBN). The experimental conclusions have demonstrated compatibility with modified insulin’s distinct cellular activity, comparably to aspart. Additionally, the hybrid structural approach combined with computational analysis employed in this study provides novel insight into the structural dynamics of newly modified and functional insulin vs insulin aspart monomeric entities. It allows further molecular understanding of intermolecular interrelations driving dissociation kinetics and, therefore, a fast action mechanism.

59 BASIC BIOLOGICAL SCIENCES↗

Structural basis for a highly conserved RNA-mediated enteroviral genome replication

Abstract Enteroviruses contain conserved RNA structures at the extreme 5′ end of their genomes that recruit essential proteins 3CD and PCBP2 to promote genome replication. However, the high-resolution structures and mechanisms of these replication-linked RNAs (REPLRs) are limited. Here, we determined the crystal structures of the coxsackievirus B3 and rhinoviruses B14 and C15 REPLRs at 1.54, 2.2 and 2.54 Å resolution, revealing a highly conserved H-type four-way junction fold with co-axially stacked sA-sD and sB-sC helices that are stabilized by a long-range A•C•U base-triple. Such conserved features observed in the crystal structures also allowed us to predict the models of several other enteroviral REPLRs using homology modeling, which generated models almost identical to the experimentally determined structures. Moreover, our structure-guided binding studies with recombinantly purified full-length human PCBP2 showed that two previously proposed binding sites, the sB-loop and 3′ spacer, reside proximally and bind a single PCBP2. Additionally, the DNA oligos complementary to the 3′ spacer, the high-affinity PCBP2 binding site, abrogated its interactions with enteroviral REPLRs, suggesting the critical roles of this single-stranded region in recruiting PCBP2 for enteroviral genome replication and illuminating the promising prospects of developing therapeutics against enteroviral infections targeting this replication platform.

Biochemistry & Molecular Biology↗

Mechanisms of Polyethylene Terephthalate Pellet Fragmentation into Nanoplastics and Assimilable Carbons by Wastewater Comamonas

Comamonadaceae bacteria are enriched on poly(ethylene terephthalate) (PET) microplastics in wastewaters and urban rivers, but the PET-degrading mechanisms remain unclear. Here, we investigated these mechanisms with Comamonas testosteroniKF-1, a wastewater isolate, by combining microscopy, spectroscopy, proteomics, protein modeling, and genetic engineering. Compared to minor dents on PET films, scanning electron microscopy revealed significant fragmentation of PET pellets, resulting in a 3.5-fold increase in the abundance of small nanoparticles (<100 nm) during 30-day cultivation. Infrared spectroscopy captured primarily hydrolytic cleavage in the fragmented pellet particles. Solution analysis further demonstrated double hydrolysis of a PET oligomer, bis(2-hydroxyethyl) terephthalate, to the bioavailable monomer terephthalate. Supplementation with acetate, a common wastewater co-substrate, promoted cell growth and PET fragmentation. Of the multiple hydrolases encoded in the genome, intracellular proteomics detected only one, which was found in both acetate-only and PET-only conditions. Homology modeling of this hydrolase structure illustrated substrate binding analogous to reported PET hydrolases, despite dissimilar sequences. Mutants lacking this hydrolase gene were incapable of PET oligomer hydrolysis and had a 21% decrease in PET fragmentation; re-insertion of the gene restored both functions. Thus, we have identified constitutive production of a key PET-degrading hydrolase in wastewater Comamonas, which could be exploited for plastic bioconversion.

54 ENVIRONMENTAL SCIENCES↗

Crystal structures of Salmonella enterica FraB deglycase reveal a conformational heterodimer with remarkable structural plasticity at the active site

Abstract Thefralocus ofSalmonella entericaencodes five genes for metabolism of fructose‐asparagine, an Amadori product formed by condensation of asparagine with glucose. In the last step of this pathway, the FraB deglycase cleaves 6‐phospho‐fructose‐aspartate into glucose‐6‐phosphate and aspartate. In homology models, FraB forms a homodimer with two equivalent active sites located at the dimer interface. E214 and H230, two invariant residues essential for catalysis, project into each active site cleft from opposing subunits of the dimer. Here, we have determined six crystal structures of FraB, three of a variant containing an N‐terminal His 6 tag and two mutations needed for crystallization (hereafter referred to as WT′), two with additional mutations to active site residues (E214A and P232A), and one of a variant with C‐terminal residues 313–325 deleted. Surprisingly, in the WT′ FraB structure, the two catalytic residues, E214 (general base) and H230 (general acid), are positioned ~22 Å apart. In the E214A and C‐terminus‐truncated FraB variants, however, a conformational change in the E214‐residing helix brings E214 and H230* to ~7 Å (* indicates residue from the second protomer that creates the inter‐subunit catalytic center). The loop bearing H230 also exhibits significant variation, ranging from being completely disordered to adopting open or closed states, with the nearby P232* residue being eithercisortrans. The C‐terminal residues 313–325 form a flexible “C‐tail” that can be fully disordered, bind in the active site to block access of substrate, or angle across the active site to wrap across the other subunit of the dimer and potentially close over substrate. Collectively, these structures reveal that FraB is a conformational heterodimer with two chemically identical subunits that are constrained to adopt different structures as they come together for catalysis. This plasticity likely involves correlated opening and closure of the two active sites for their respective binding and release of substrates and ligands.

Biochemistry & Molecular Biology↗

Improved Biofuel Production through Discovery and Engineering of Terpene Metabolism in Switchgrass

Project Objectives - Of the myriad specialized metabolites that plants deploy to adapt to environmental challenges, terpenes form the largest group. In many major crops, unique terpene blends serve as key stress defenses that directly impact plant fitness and yield. In addition, terpenes, such as bisabolene and pinene, are used for producing renewable biofuels. Essential to advancing a broader use of terpenes for biofuel feedstock engineering is a system-wide knowledge of the diverse biosynthetic machinery and defensive potential of often species-specific terpene blends. The proposed project would merge genome-wide enzyme discovery with comparative –omics, protein structural and plant microbiome studies to define the biosynthesis and stress-defensive functions of the switchgrass (Panicum virgatum) terpene network. These insights would be combined with developing and applying non-transgenic genome editing tools to design plants with desirable terpene blends for higher productivity and biofuel production on marginal lands. As a dedicated lignocellulosic feedstock for U.S. biofuel production with high net energy yield, stress tolerance, and available genome resources, switchgrass is well-suited for devising new avenues for biofuel production. Project Description – The diversity of plant terpene defenses is governed by species-specific families of terpene synthase (TPS) and cytochrome P450 monooxygenase (P450) enzymes. Mining of the switchgrass genome (genotype Alamo) identified ~100 TPS and P450 candidate genes, and combinatorial biochemical analysis of synthesized TPSs and P450s revealed more than a dozen enzymes with common and novel activities. In addition, several identified terpene metabolites and the corresponding transcripts were up-regulated in response to abiotic stressors. These findings demonstrate a unique switchgrass terpene network with probable importance to abiotic stress tolerance, thus providing a large chemical portfolio for optimizing crop resistance, yield, and biofuel composition. Leveraging these preliminary data, we propose to generate a genome-wide map of the switchgrass terpene metabolic network through multi-gene co-expression analyses that allow the efficient cross-validation of TPS and P450 functions. Key enzymes would further be applied to structure-function studies via X-ray protein crystallography, homology modeling and site-directed mutagenesis to gain mechanistic insight into the catalytic specificity of switchgrass terpene metabolism and provide gene and amino acid targets for genome editing. In tandem with terpene pathway discovery, system-wide metabolomics, transcriptomics and proteomics studies in switchgrass accessions of contrasting drought tolerance would define the role of switchgrass terpene metabolism in conferring abiotic stress resilience. Metabolic changes would be assessed in a combined approach of targeted (terpenes) and untargeted metabolite profiling using a high-resolution LC-MS/MS approach, differential gene expression analyses through multiplexed Illumina RNA sequencing, and quantitative analysis of high-priority pathway enzymes using multiple reaction monitoring (MRM). Drawing on these insights, knock-down/out mutants of stress-associated pathway nodes would be generated by optimizing transient virus-induced gene silencing (VIGS) and CRISPR/Cas9 systems under control of the Tobacco Rattle Virus (TRV). The resulting mutant lines would then be analyzed for stress susceptibility and the impact on the root microbiome to define gene functions in planta. Knowledge of terpene pathways, enzyme mechanisms and bioactivities would be applied to enhance switchgrass stress resilience and to tailor-make terpene blends for biofuel production. Here, TRV-enabled CRISPR/Cas9 genome editing, including allele-specific knock-out of redundant genes, engineering of enzyme specificity via structure-guided point mutations, and overexpression of terpene genes relevant to stress-protection or biofuel production, would be used to increase metabolic flux toward desired pathways. Broader Impacts - Integrating the system-wide discovery, mechanistic analysis and non-transgenic genome engineering of the switchgrass terpene network aligns the required steps to unlock the chemical potential of this important metabolite class to generate crops that are more resistant to stress and provide advanced biofuel production in light of rising climate pressures as foreseeable challenges for bioenergy crop cultivation. The proposed project would further offer interdisciplinary student training through active involvement in the project and integration of research concepts and outcomes into newly-developed graduate and undergraduate courses on Plant Biotechnology.

09 BIOMASS FUELS↗

Cannabis monoterpene synthases: evaluating structure–function relationships

Terpene synthases catalyze the first committed step in the biosynthesis of terpenes, a structurally diverse class of natural products that also encompasses volatiles derived from precursors in the C10 to C15 range (termed monoterpenes and sesquiterpenes, respectively). In the review section of this article, we are providing information about all functionally characterized monoterpene synthases (MTSs) and sesquiterpene synthases (STSs) of Cannabis sativa L. We are also exploring the locations of MTSs and STSs in the chromosome-level assembly of the reference chemovar CBDRx. A follow-up computational structure–function analysis focuses on MTSs, as there is already a rich literature available on the topic. More specifically, by employing sequence comparisons and homology structural modeling, we infer which amino acid residues are likely to constrain the available space in the active site of cannabis MTSs. The emphasis of these studies was to investigate why some MTSs accept only a C10 diphosphate as substrate, while mixed MTS/STS enzymes also accommodate a C15 diphosphate. Here, by combining a literature review and computational analyses in a hybrid format, we are laying the foundation for future studies to better understand the determinants of substrate and product specificity in these fascinating enzymes.

59 BASIC BIOLOGICAL SCIENCES↗

Osteocalcin binds to a GPRC6A Venus fly trap allosteric site to positively modulate GPRC6A signaling

GPRC6A, a member of the Family C G-protein coupled receptors, regulates energy metabolism and sex hormone production and is activated by diverse ligands, including cations, L-amino acids, the osteocalcin (Ocn) peptide and the steroid hormone testosterone. We sought a structural framework for the ability of multiple distinct classes of ligands to active GPRC6A. We created a structural model of GPRC6A using Alphafold2. Using this model we explored a putative orthosteric ligand binding site in the bilobed Venus fly trap (VFT) domain of GPRC6A and two positive allosteric modulator (PAM) sites, one in the VFT and the other in the 7 transmembrane (7TM) domain. We provide evidence that Ocn peptides act as a PAM for GPRC6A by binding to a site in the VFT that is distinct from the orthosteric site for calcium and L-amino acids. In agreement with this prediction, alternatively spliced GPRC6A isoforms 2 and 3, which lack regions of the VFT, and mutations in the computationally predicted Ocn binding site, K352E and H355P, prevent Ocn activation of GPRC6A. These observations explain how dissimilar ligands activate GPRC6A and set the stage to develop novel molecules to activate and inhibit this previously poorly understood receptor.

59 BASIC BIOLOGICAL SCIENCES↗

A hybrid surrogate modeling framework for the Digital Twin of a Fluoride-salt-cooled High-temperature Reactor (FHR)

While nuclear energy is a non-greenhouse-gas emitting energy source, expensive operational costs due to the high-level of safety requirements decreases their competitiveness in the sustainable energy market. Advanced reactor concepts paired with Digital Twins aim to increase the commercialization gains of nuclear energy by reducing operational costs, increasing reactor reliability and enhancing power generation. To support Digital Twin tasks such as real-time autonomous control, proactive maintenance monitoring or optimizing power demand operations, a fast and accurate virtual representation of the Nuclear Power Plant (NPP) is required. The computational cost of high-fidelity, physics-based models are unsuitable for real-time analysis or scalability. Here, in this work, a hybrid surrogate modeling framework is developed fora Fluoride-salt-cooled High-temperature Reactor (FHR) that leverages physics-inspired models for key reactor components and uses data-driven methods for rapid system state space prediction. The Xenon reactivity feedback model is integrated to inform the surrogate model about the reactor core and the homologous pump theory model is the basis for representing pump degradation. Using a detailed, two dimensional thermal hydraulics model to generate data on the FHR, we train a network of Vectorized Autoregressive Moving-Average with eXogenous input (VARMAX) models to predict the remaining state values. The result is a surrogate model that provides a detailed reactor state representation of 41 system states and a pump degradation analysis. The framework is applied to Load Follows profiles, yielding high accuracy and a speedup that is more than 4000x faster compared to the higher- fidelity thermal hydraulics model, enabling real-time operational intelligence and applications in long horizon predictions. While the surrogate model framework is demonstrated for the particular case of FHR, the hybrid physical/data-driven modeling approach including the network of surrogates and the underlying modularity has the potential to be applied to other physical asset systems.

Digital Twins↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Structural determination of a full-length plant cellulose synthase informed by experimental and in silico methods

Three-dimensional structure determination and prediction of proteins with intrinsically disordered regions, unstructured regions, conformational flexibility, and lacking homologous structures are challenging. We previously predicted and refined an in silico structure of a plant cellulose synthase from cotton (GhCESA1), and more recently, cryo-electron microscopy (cryo-EM) has resolved a majority of the lengths of two CESA structures from poplar (PttCESA8) and cotton (GhCESA7). However, 26–30% of these cryo-EM structures remain unresolved, including the N-terminal domain, half of the class-specific region, the gating loop region, and the C-terminal domain. Here, we describe the generation and evaluation of a full-length hybrid GhCESA1 model based on this cryo-EM PttCESA8 structure, with unresolved regions completed using this in silico refined GhCESA1 model. All-atom molecular dynamics simulations and subsequent energy minimizations were performed for the in silico and hybrid GhCESA1 models in a lipid bilayer-water-ion environment, and structural stability, dynamics, energetics, contacts, and quality were evaluated. The unresolved regions were found to be the most dynamic, in agreement with their poor electron density with cryo-EM. The hybrid model exhibited a higher total secondary structure content, more favorable intra-protein and protein-lipid interaction energies, and improved quality metrics. Moreover, hydrogen bonding was revealed to be a primary mechanism for intra-protein and protein-lipid contacts. These results demonstrate that in silico structure prediction and refinement may be useful to augment experimental structure determination, especially for disordered and unstructured regions. Furthermore, this hybrid model can serve as a steppingstone to derive full-length homology models of other CESAs found in more experimentally tractable organisms.

59 BASIC BIOLOGICAL SCIENCES↗

Empirical evidence that glucan-interacting amino acid side chains within the transmembrane channel collectively facilitate cellulose synthase function

The fundamental mechanism of cellulose synthesis is widely conserved across Kingdoms and depends on cellulose synthases, which are processive, dual-function, family 2 glycosyltransferases (GT-2). These enzymes polymerize glucose on the cytoplasmic side of the plasma membrane and export the glucan chain to the cell surface through an integral transmembrane (TM) channel. Structural studies of active plant cellulose synthases (CESAs) have revealed interactions between the nascent glucan chain and the side chains of polar, charged, and aromatic amino acid residues that line the TM channel. However, the functional consequences of modifying these side chains have not been tested in vivo in CESAs or other processive GT-2s. To test this, we used an established in vivo assay based on genetic complementation of CESA5 in the moss, Physcomitrium patens. For accurate prediction of glucan-interacting amino acid residues, we generated a complete homotrimeric molecular model of PpCESA5 using a combination of homology and de novo modeling. All-atom molecular dynamics-based analyses of contact metrics and interaction energy identified 23 amino acid residues with high propensity to interact with the nascent glucan chain within the TM channel or on the apoplastic surface of PpCESA5. Mutating any one of 18 of these amino acid residues to alanine, thereby removing their side chains, abolished or impaired CESA function, with the strongest effects observed upon the loss of charged amino acid side chains. This provides direct evidence to support the hypothesis that multiple amino acid residues collectively maintain a smooth energy landscape within the TM channel to facilitate glucan translocation.

59 BASIC BIOLOGICAL SCIENCES↗

A Wox3 -patterning module organizes planar growth in grass leaves and ligules

Grass leaves develop from a ring of primordial initial cells within the periphery of the shoot apical meristem, a pool of organogenic stem cells that generates all of the organs of the plant shoot. At maturity, the grass leaf is a flattened, strap-like organ comprising a proximal supportive sheath surrounding the stem and a distal photosynthetic blade. The sheath and blade are partitioned by a hinge-like auricle and the ligule, a fringe of epidermally derived tissue that grows from the adaxial (top) leaf surface. Together, the ligule and auricle comprise morphological novelties that are specific to grass leaves. Understanding how the planar outgrowth of grass leaves and their adjoining ligules is genetically controlled can yield insight into their evolutionary origins. Here we use single-cell RNA-sequencing analyses to identify a ‘rim’ cell type present at the margins of maize leaf primordia. Cells in the leaf rim have a distinctive identity and share transcriptional signatures with proliferating ligule cells, suggesting that a shared developmental genetic programme patterns both leaves and ligules. Moreover, we show that rim function is regulated by genetically redundant Wuschel-like homeobox3 (WOX3) transcription factors. Higher-order mutations in maize Wox3 genes greatly reduce leaf width and disrupt ligule outgrowth and patterning. Together, these findings illustrate the generalizable use of a rim domain during planar growth of maize leaves and ligules, and suggest a parsimonious model for the homology of the grass ligule as a distal extension of the leaf sheath margin.

59 BASIC BIOLOGICAL SCIENCES↗

A multi-ancestry GWAS of Fuchs corneal dystrophy highlights the contributions of laminins, collagen, and endothelial cell regulation

Fuchs endothelial corneal dystrophy (FECD) is a leading indication for corneal transplantation, but its molecular etiology remains poorly understood. We performed genome-wide association studies (GWAS) of FECD in the Million Veteran Program followed by multi-ancestry meta-analysis with the previous largest FECD GWAS, for a total of 3970 cases and 333,794 controls. We confirm the previous four loci, and identify eight novel loci: SSBP3, THSD7A, LAMB1, PIDD1, RORA, HS3ST3B1, LAMA5, and COL18A1. We further confirm the TCF4 locus in GWAS for admixed African and Hispanic/Latino ancestries and show an enrichment of European-ancestry haplotypes at TCF4 in FECD cases. Among the novel associations are low frequency missense variants in laminin genes LAMA5 and LAMB1 which, together with previously reported LAMC1, form laminin-511 (LM511). AlphaFold 2 protein modeling, validated through homology, suggests that mutations at LAMA5 and LAMB1 may destabilize LM511 by altering inter-domain interactions or extracellular matrix binding. Finally, phenome-wide association scans and colocalization analyses suggest that the TCF4 CTG18.1 trinucleotide repeat expansion leads to dysregulation of ion transport in the corneal endothelium and has pleiotropic effects on renal function.

59 BASIC BIOLOGICAL SCIENCES↗

Topological Signatures of Adversaries in Multimodal Alignments

Topological Data Analysis for Adversarial Detection (LANL O4937) - Detects adversarial examples in vision-language models using persistent homology and two-sample testing. Combines TDA features from CLIP embeddings with statistical methods (ME, SCF, SAMMD, C2ST) for robust detection across ImageNet, CIFAR-10/100.

Bhattarai, Manish↗

Bacterial hemophilin homologs and their specific type eleven secretor proteins have conserved roles in heme capture and are diversifying as a family

Cellular life relies on enzymes that require metals, which must be acquired from extracellular sources. Bacteria utilize surface and secreted proteins to acquire such valuable nutrients from their environment. These include the cargo proteins of the type eleven secretion system (T11SS), which have been connected to host specificity, metal homeostasis, and nutritional immunity evasion. This Sec-dependent, Gram-negative secretion system is encoded by organisms throughout the phylum Proteobacteria, including human pathogens Neisseria meningitidis, Proteus mirabilis, Acinetobacter baumannii, and Haemophilus influenzae. Experimentally verified T11SS-dependent cargo include transferrin-binding protein B (TbpB), the hemophilin homologs heme receptor protein C (HrpC), hemophilin A (HphA), the immune evasion protein factor-H binding protein (fHbp), and the host symbiosis factor nematode intestinal localization protein C (NilC). Here, we examined the specificity of T11SS systems for their cognate cargo proteins using taxonomically distributed homolog pairs of T11SS and hemophilin cargo and explored the ligand binding ability of those hemophilin cargo homologs. In vivo expression in Escherichia coli of hemophilin homologs revealed that each is secreted in a specific manner by its cognate T11SS protein. Sequence analysis and structural modeling suggest that all hemophilin homologs share an N-terminal ligand-binding domain with the same topology as the ligand-binding domains of the Haemophilus haemolyticus heme binding protein (Hpl) and HphA. We term this signature feature of this group of proteins the hemophilin ligand-binding domain. Network analysis of hemophilin homologs revealed five subclusters and representatives from four of these showed variable heme-binding activities, which, combined with sequence-structure variation, suggests that hemophilins are diversifying in function.

59 BASIC BIOLOGICAL SCIENCES↗

Quantitative SANS and multi-model analysis of spacer-dependent micellization of urea-based gemini surfactants

The micellization behavior of urea-based cationic gemini surfactants was investigated using small-angle neutron scattering (SANS) with multi-model form factor analysis. A homologous series of surfactants with urea group included in the hydrophobic tail and polymethylene spacers consisting of two to ten methylene units was analyzed using three form factor models: a core–shell ellipsoid and two variants of homogeneous ellipsoids. The results from all models show a consistent trend of the micelle structures, confirming that the spacer length critically influences micellar geometry, aggregation number, and hydration. The surfactant with four CH 2 groups in the spacer formed the largest micelles with the highest aggregation number, while longer spacers led to progressively smaller, more compact aggregates. The shell hydration—quantified as the volume fraction of heavy water within the hydrophilic region—decreased systematically with increasing spacer length due to enhanced hydrophobicity of the headgroup-spacer region. Intermicellar interactions, modeled as screened Coulomb interaction using the rescaled mean spherical approximation (RMSA), revealed the strongest electrostatic repulsion for the case of four methylene groups in the spacer, corresponding to the highest micellar charge and largest interparticle spacing. The observed spacer-dependent trends were robust across all modeling approaches, demonstrating that the spacer length serves as a key structural determinant of self-assembly in this type of urea-based gemini systems. These findings provide insight into the design of gemini surfactants with tailored aggregation behavior for applications in drug delivery, nanostructure templating, and solubilization technologies.

Core–shell ellipsoid model↗

Harnessing Machine Learning and Data Fusion for Accurate Undocumented Well Identification in Satellite Images

This study utilizes satellite data to detect undocumented oil and gas wells, which pose significant environmental concerns, including greenhouse gas emissions. Three key findings emerge from the study. Firstly, the problem of imbalanced data is addressed by recommending oversampling techniques like Rotation–GaussianBlur–Solarization data augmentation (RGS), the Synthetic Minority Over-Sampling Technique (SMOTE), or ADASYN (an extension of SMOTE) over undersampling techniques. The performance of borderline SMOTE is less effective than that of the rest of the oversampling techniques, as its performance relies heavily on the quality and distribution of data near the decision boundary. Secondly, incorporating pre-trained models trained on large-scale datasets enhances the models’ generalization ability, with models trained on one county’s dataset demonstrating high overall accuracy, recall, and F1 scores that can be extended to other areas. This transferability of models allows for wider application. Lastly, including persistent homology (PH) as an additional input improves performance for in-distribution testing but may affect the model’s generalization for out-of-distribution testing. A careful consideration of PH’s impact on overall performance and generalizability is recommended. Overall, this study provides a robust approach to identifying undocumented oil and gas wells, contributing to the acceleration of a net-zero economy and supporting environmental sustainability efforts.

SMOTE↗