Search NASA⌕ Search

SEARCH · Search NASA

Results for “web”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Effects of dissolved cations, dissolved organic carbon, and exposure concentrations on per- and polyfluoroalkyl substances bioaccumulation in freshwater algae

Per- and polyfluoroalkyl substances (PFAS) have attracted global attention because of their persistence, toxicity, bioaccumulation potential, and associated adverse effects. As important primary producers, freshwater algae constitute the base of the food web in freshwater aquatic ecosystems. However, the effects of key environmental factors on PFAS uptake and bioaccumulation in freshwater algae have not been thoroughly studied. In this study, three bioaccumulation experiments were conducted to evaluate the influence of dissolved cations, dissolved organic carbon, and exposure concentrations on PFAS bioaccumulation in algae. Among the 14 studied PFAS, seven long-chain PFAS tended to bioaccumulate in algae. Elevated divalent cations (Ca 2+ and Mg 2+ ) and dissolved organic carbon did not significantly change the algal bioconcentration factors (BCFs) of PFAS, suggesting complexity of the interactions among PFAS, environmental factors, and biotic activities. Additionally, increasing exposure concentrations (0.5, 1, 5, and 10 μg/L of each PFAS) increased PFAS concentrations in algae but decreased the BCF values. This indicated that attention should be paid to the application of BCFs in future studies, including ecological risk assessment. Moreover, fluorotelomer sulfonic acids (FTSs) were incompletely recovered, suggesting that biotransformation occurred. Further studies should be conducted to evaluate whether algae play a role in FTSs biotransformation and to determine the mechanisms. In conclusion, studying the impacts of key environmental factors on PFAS bioaccumulation in algae is crucial for understanding the bioaccumulation processes that occur at the lowest trophic level and that eventually affect the dynamics of entire aquatic ecosystems.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗

Process intensification approach to enhancing heat and mass transfer during drying: Ultrasonic (US) assisted drying of paper and board

Drying of paper and board is conventionally achieved through alternating conduction (steam-heated cylinders) and pocket convection (heated air over the paper web surface). These conventional drying systems rely heavily on steam from fossil fuels, resulting in inefficiencies, high energy usage, and thermal losses due to surface-driven mechanisms. Here, to address these challenges, an experimental system, with in-situ drying characteristics measurements, was developed to investigate process intensification using ultrasonic-based dewatering—a volumetric, pressure-driven acoustic energy system—integrated with conventional drying. The objectives of this study are to assess the impact of ultrasonics (US) on dewatering; compare performances to conventional drying systems; identify improvements in drying rate and energy use as a function of moisture content; and gain potential insights on heat and mass transfer mechanisms during US-assisted drying. US performance was evaluated across frequencies, power levels, pulp types, and basis weights. Results show that improvements to ultrasonic applications in conjunction with convection were 30-43% in drying rate and 20-35% in drying time over continuous and intermittent applications. When combined with conduction and convection, ultrasonics yielded up to 20% improvement in both rate and time and up to 20% reduction in energy consumption. Observations support a hypothesis of extension of the constant rate period due to improved capillary flow at higher moisture content and enhancing vapor diffusion and boundary layer disruption at lower moisture contents during falling rate period. These findings will inform future modeling, simulation, design and optimization of advanced drying systems.

42 ENGINEERING↗

Process intensification approach to enhancing heat and mass transfer: Radio frequency (RF) assisted drying of paper and board

Conventional multi-cylinder drying of paper and board typically relies on both conductive drying from steam-heated cylinders and convective drying, where heated air flows over the paper web surface. Conduction primarily contributes to heat transfer, while convection is the main driver of mass transfer. However, conventional drying systems are heavily dependent on steam, usually powered by fossil fuels, and are often energy-inefficient with high levels of waste. Additionally, these surface-driven processes result in a lower percentage of energy absorption compared to the energy supplied, leading to significant energy losses. To improve this long-standing process, an experimental system was developed to investigate a process intensification approach involving the integration of Radio Frequency (RF) heating, a volumetric electromagnetic technology, alongside traditional conduction and convection drying methods. This study also emphasizes the use of sensors to continuously monitor key parameters such as moisture content, supply system temperatures, sample temperatures, air flows, and drying rates in real-time. The effect of RF as an auxiliary energy source in localized environments at varying moisture levels was explored to optimize industrial drying systems, quantify potential improvements, and provide insights for future studies. Experimental results from trials combining RF with convection and with the base case alternating conduction-convection drying processes are presented. It is shown that RF is a viable process intensification approach for paper drying improving the drying rate and energy intensity at higher moisture contents. These findings offer valuable insights for process intensification and contribute to the process development, modeling, and simulation of advanced paper drying techniques.

42 ENGINEERING↗

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Synergistic cellulase–xylanase formulations for enhanced dewatering and fiber bonding toward energy-efficient and sustainable paper and packaging production

A mechanistic understanding of the synergistic effects of enzymes on cellulosic fibers dewatering and fiber bonding is essential for advancing energy-efficiency and lightweight production of paper and packaging materials. This study investigates the impact of varying cellulase and xylanase formulations on equilibrium moisture content (EMC) after pressing and tensile strength of cellulosic fiber webs using a factorial experimental design. Nine custom enzyme formulations were evaluated at controlled dosages ranging from ∼20 to 76 FPU/mL cellulase and ∼500–1130 IU/mL xylanase. Response surface modeling revealed a significant synergy, particularly at 20–40 FPU/mL cellulase combined with ≥ 1000 IU/mL xylanase. Under these conditions, EMC decreased by up to 3.6% compared with the untreated refined control, while tensile index gains of up to 16% were statistically significant for optimized blends (p < 0.05). The interaction between cellulase and xylanase was also significant for the tensile response (p = 0.010). Protein efficiency analysis showed that optimized formulations containing 30–40% less protein outperformed the commercial benchmark. The nonlinear synergy between cellulase and xylanase is attributed to their complementary substrate specificities. Endoglucanase- and β‑glucosidase‑rich cellulases hydrolyze internal β‑1,4‑glycosidic bonds in amorphous cellulose, loosening fiber walls and increasing flexibility, while xylanases target hemicellulose, primarily xylan-rich domains, enhancing porosity and improve cellulase accessibility. Tailoring enzyme formulations at low loadings overcomes traditional trade-offs between strength and dewatering, enabling cost-effective, energy-efficient, low-carbon solutions for sustainable packaging and hygiene products.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Scintillator Library: A database of inorganic and organic scintillator properties

The Scintillator Library (scintillator.lbl.gov) is a database of scintillator properties hosted by Lawrence Berkeley National Laboratory in a web-accessible format. It contains a variety of measured inorganic and organic scintillator properties extracted from peer-reviewed literature and manufacturer specifications. Data housed within the Scintillator Library supply an important resource for developers of scintillator-based detection systems and an aid for scientists seeking to establish connections between fundamental material and chemical properties and the associated scintillation performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A methodology for decay heat characterization in molten salt reactors

Accurate decay heat prediction in molten salt reactors (MSRs) faces dual challenges: complex operational uncertainties and the need for interpretable models compatible with engineering workflows. This work presents a hybrid machine learning and segmented polynomial methodology that addresses both requirements through three key innovations. First, a modular data architecture encodes MSR-specific operational parameters (power density: 1-100 W cm -3 , humidity: 0-0.1 wt %, air ingress: 0-0.1 mol %) with uncertainty-aware temporal discretization spanning 15 orders of magnitude. Second, region-optimized machine learning models achieve 92.3 % root mean square error (RMSE) reduction over conventional polynomials while maintaining physical interpretability through automated piecewise equation generation. Third, dual front-end interfaces accelerate safety analyses — a Jupyter environment enables researchers to explore 10,000+ parameter combinations via interactive widgets, while a Streamlit web application reduces design iteration cycles through production-grade visualization tools. Operational deployment demonstrates prediction times of only a couple hundred milliseconds for 10 4 years decay profiles, enabling real-time optimization of spent fuel container designs.

42 - ENGINEERING↗

Alchemy: A Model-Based Approach for 2D to 3D Autonomous Nuclear System Design

Engineering design of nuclear power plant (NPP) piping and equipment systems frequently bypasses crucial 2D system planning, instead moving straight to 3D modeling. This often leads to designs that exceed building envelope constraints, forcing expensive and time-consuming redesigns. When 2D modeling is employed, it typically involves labor-intensive manual workflows that convert 2D drawings into 3D models, resulting in inefficiencies and errors across design iterations. These workflows further suffer from poor software interoperability and dependence on proprietary software ecosystems, thereby contributing to schedule delays and cost overruns. This paper presents Alchemy, an autonomous framework that transforms 2D system definitions into Industry Foundation Classes (IFC)-compliant 3D building information models (BIMs) for expediting nuclear facility design at the conceptual preliminary phase. Using a model-based approach, the framework treats the 2D system diagram as the central reference model employed to automatically generate all subsequent outputs, ensuring consistency between the system definition and the resulting physical design. A web-based interface enables engineers to define hierarchical system topologies including associated equipment, geometric properties, and connectivity requirements. A two-phase equipment layout optimization algorithm automatically computes collision-free spatial configurations within predefined building envelopes. An artificial intelligence (AI)-assisted pipe routing module then generates orthogonal, collision-free routing paths, allowing the user to select either an A* search-based method or an Ant Colony Optimization (ACO)-based method. All outputs are authored natively in IFC format, relying on open-source technologies and standardized formats in order to ensure extensibility and eliminate proprietary software dependencies. The proposed framework is validated on two representative pressurized-water reactor (PWR)-based case studies, for which it autonomously generates IFC-compliant 3D models in minutes, drastically reducing workflows that typically require hours of manual effort. The generated model demonstrates topologically correct equipment placement, physically plausible spatial relationships, and collision-free pipe routing consistent with known PWR loop configurations. This work represents a foundational step toward digital engineering for nuclear facility preliminary design, with future ongoing development targeting design code compliance and expanded system complexity.

97 - MATHEMATICS AND COMPUTING↗

Mapping heat vulnerability in cities: A tale of two california cities

Extreme heat is a major cause of weather-related deaths in the United States. To address this, a heat vulnerability index (HVI) is crucial for assessing heat risk and identifying vulnerable urban areas and populations, supporting city planning and emergency response. Current HVI studies often use Principal Component Analysis (PCA) on environmental, socioeconomic, and medical data to aggregate vulnerability indicators into a single index. However, these fixed aggregation weights struggle to adapt to different use cases, which may require varying focuses. Moreover, existing tools primarily consider outdoor heat exposure, providing an incomplete picture of actual exposure, as people spend most of their time indoors. Our research introduces an HVI web mapping tool that addresses these gaps in the literature by: (1) allowing flexible weights to adapt to different use cases, and (2) uniquely integrating both outdoor and indoor heat exposure by considering building characteristics for a more comprehensive risk assessment. We demonstrated this tool in two California cities with contrasting climates: Fresno (inland, arid, hot summers) and Oakland (temperate coastal). This HVI mapping tool provides essential decision support for policymakers and stakeholders in both short-term heat mitigation and long-term urban planning for building interventions and infrastructure development.

BES↗

CHARMM-GUI Bicelle Builder : An Extension of Membrane Builder for Modeling and Simulation of Bicelle Systems

Membrane mimetics, such as detergent micelles, nanodiscs, and amphipol complexes, which can provide membrane-like environments while retaining small and soluble features, have been utilized to study membrane proteins. A bicelle, composed of varying lipids and detergents, is a useful membrane mimetic because the lipid-to-detergent ratio, the q-value, can be adjusted to alter the properties of the aggregate, including the thickness and size of the bicelle. However, building a bicelle model for modeling and simulation studies requires nontrivial efforts, even for experts. We introduce CHARMM-GUI Bicelle Builder, a web-based platform that can generate various all-atom bicelle systems via a graphical user interface with all available lipids and detergents in Membrane Builder. To illustrate and validate Bicelle Builder with practical systems, we have modeled and simulated pure bicelles consisting of 1,2-dimyristoyl-sn-glycero-3-phosphocholine (DMPC) lipids with 1,2-dihexanoyl-sn-glycero-3-phosphocholine (C6DHPC) detergents and protein–bicelle complexes, composed of DMPC with C6DHPC, foscholine-10 (FOS10), and lysophosphatidylcholine-12 (LPC12) detergents. Our simulation results indicate that Bicelle Builder can generate reliable and robust bicelle models with and without proteins that retain DMPC bilayer characteristics. Bicelle Builder is expected to help researchers better understand not only bicelles themselves but also atomistic-level structures of protein–bicelle complexes that are often difficult to access through experimental approaches.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Models and Algorithms for Equilibrium Analysis of Mixed-Material Nucleic Acid Systems

Dynamic programming algorithms within the NUPACK software suite enable analysis of equilibrium base-pairing properties for complex and test tube ensembles containing arbitrary numbers of interacting nucleic acid strands. Currently, calculations are limited to single-material systems that are either all-RNA or all-DNA. Here, to enable analysis of mixed-material systems that are critical for modern applications in vitro, in situ, and in vivo, we develop physical models and dynamic programming algorithms that allow the material of the system to be specified at nucleotide resolution. Free energy parameter sets are constructed for both RNA/DNA and RNA/2'OMe-RNA mixed-material systems by combining available empirical mixed-material parameters with single-material parameter sets to enable treatment of the full complex and test tube ensembles. New dynamic programming recursions account for the material of each nucleotide throughout the recursive process. For a complex with N nucleotides, the mixed-material dynamic programming algorithms maintain the O(N 3 ) time complexity of the single-material algorithms, enabling efficient calculation of diverse physical quantities over complex and test tube ensembles (e.g., complex partition function, equilibrium complex concentrations, equilibrium base-pairing probabilities, minimum free energy secondary structure(s), and Boltzmann-sampled secondary structures) at a cost increase of roughly 2.0-3.5×. The results of existing single-material algorithms are exactly reproduced when applying the new mixed-material algorithms to single-material systems. Accuracy is significantly enhanced using mixed-material models and algorithms to predict RNA/DNA and RNA/2'OMe-RNA duplex melting temperatures from the experimental literature as well as RNA/DNA melt profiles from new experiments. In conclusion, mixed-material analyses can be performed online using the NUPACK web app (www.nupack.org) or locally using the NUPACK Python module.

2′OMe-RNA↗

Vision and Development of a Design, Implementation, and Verification Automation (DIVA) Software Platform for DNA Construction

Abstract DNA construction, while a prerequisite to many biological endeavors, is often a time-consuming distraction from an individual’s primary research objectives. We envisioned that with the right software infrastructure and cultural mindset, a single person could execute in parallel the batched DNA construction tasks of an entire research institute, at scales realizing efficiency gains through process and laboratory automation. In pursuit of this vision, we developed the Design, Implementation, and Verification Automation (DIVA) software platform. DIVA’s web interface enables researchers to design DNA constructs (using visual biological computer-aided design tools and biological parts repositories), submit designs for construction to dedicated staff, and track DNA construction as it progresses. DIVA supports the dedicated staff through the DNA construction process and records both successful and unsuccessful attempts toward improving the overall process. The platform is publicly available at public-diva.jbei.org and its open-source code through github.com/JBEI/DIVA.

Plahar, Hector [DOE Agile BioFoundry , , ,; DOE Jo↗

Expanded Understanding of the Western Antarctic Peninsula Sea‐Ice Environment Through Local and Regional Observations at Palmer Station

Abstract The Western Antarctic Peninsula (WAP) has been experiencing rapid regional warming since at least the 1950s, however, the impacts of this warming at the local scale are variable and nuanced. Previous studies that have linked sea‐ice variability to biogeochemical cycles and food web dynamics often combine local‐scale biogeochemical data with coarse‐resolution regional satellite sea‐ice data, which may not adequately capture local sea‐ice conditions. In this study, we analyzed local‐scale in situ sea‐ice observations collected as part of a 28‐year record (1992–2020) from the Palmer Long‐Term Ecological Research site at Anvers Island, mid‐WAP, in conjunction with isotopically‐derived sea‐ice meltwater (SIM) fractions and satellite‐derived sea‐ice motion and concentration, to quantify the variability and long‐term trends in local sea‐ice behavior. In situ sea ice observations at Palmer Station displayed higher variability than satellite observations and showed no significant declines over this time, despite region‐wide declines identified in prior studies. Higher spring SIM fractions were attributed to strong northward sea‐ice motion throughout the winter. Applying these local‐scale sea‐ice insights to similarly scaled stratification and chlorophyll‐ a measurements, we found that a longer‐lasting, more consistent sea‐ice pack led to greater water column stratification following the spring sea‐ice retreat. Greater sea‐ice persistence and stronger stratification led to larger peaks in chlorophyll‐ a , though sea‐ice metrics did not explain the positive temporal trends in either stratification strength or chlorophyll‐ a . Through this study, we identify how local sea‐ice observations and meltwater data can enhance satellite data to build an understanding of the intricate connections between ice, water column dynamics, and phytoplankton.

Goodell, E.↗

The structural basis for 2′−5′/3′−5′-cGAMP synthesis by cGAS

Abstract cGAS activates innate immune responses against cytosolic double-stranded DNA. Here, by determining crystal structures of cGAS at various reaction stages, we report a unifying catalytic mechanism. apo-cGAS assumes an array of inactive conformations and binds NTPs nonproductively. Dimerization-coupled double-stranded DNA-binding then affixes the active site into a rigid lock for productive metal•substrate binding. A web-like network of protein•NTP, intra-NTP, and inter-NTP interactions ensures the stepwise synthesis of 2′−5′/3′−5′-linked cGAMP while discriminating against noncognate NTPs and off-pathway intermediates. One divalent metal is sufficient for productive substrate binding, and capturing the second divalent metal is tightly coupled to nucleotide and linkage specificities, a process which manganese is preferred over magnesium by 100-fold. Additionally, we elucidate how mouse cGAS achieves more stringent NTP and linkage specificities than human cGAS. Together, our results reveal that an adaptable, yet precise lock-and-key-like mechanism underpins cGAS catalysis.

59 BASIC BIOLOGICAL SCIENCES↗

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗