Search NASA⌕ Search

SEARCH · Search NASA

Results for “open datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Measurement of 𝐵 meson production fraction ratios in proton-proton collisions at √𝑠 = 13 TeV using open-charm and charmonium decays

Production fraction ratios of 𝐵 + , 𝐵 0 , and 𝐵$^{0}_{𝑠}$ mesons are measured in proton-proton collisions at √𝑠 = 13 TeV using a special dataset recorded in 2018 with high-rate triggers designed to collect an unbiased sample of 1⁢010 𝑏 hadrons with the CMS experiment at the LHC. These data allow the study of the open-charm decays of 𝐵 mesons (𝐵 (𝑠) → 𝜋⁢𝐷 (𝑠) ) where the 𝐷 meson decays into fully hadronic final states. By utilizing known branching fractions and precise theoretical calculations, production fraction ratios as functions of 𝐵 meson transverse momentum (𝑝 T ) and rapidity (𝑦) are measured using open-charm decays in the kinematic range of 8 < 𝑝 T < 60 GeV and |𝑦| <2.25. In addition, the same dataset is used to measure the relative production fraction ratios with the charmonium decay channels [𝐵 (𝑠) →𝑋⁢𝐽/𝜓 with 𝑋 indicating a 𝐾 + , 𝐾*(892) 0 , or 𝜙⁡(1020) meson] where the 𝐽/𝜓 meson decays into a pair of muons. The open-charm results are used to normalize the relative production fraction ratios obtained from the charmonium samples. Measurements of the ratios of branching fractions of 𝐵 meson decays to charmonium and open-charm final states are also reported, which will improve the world-average values of these ratios. Finally, we test isospin invariance in 𝐵 meson production in proton-proton collisions and observe that it holds within the experimental precision.

Bottom mesons↗

Upper-air soundings collected during the CROCUS Urban Canyons 2024 campaign in Chicago, Illinois USA

Funded by the Department of Energy’s Office of Science, Biological and Environmental Research program, Community Research on Climate and Urban Science (CROCUS) studies urban climate change and the impact it has on communities, with particular focus on disinvested, under-resourced communities. This information leads to new insights on urban climate challenges and informs future actions for mitigating and adapting to climate change at the street, neighborhood and regional levels.As part of the CROCUS effort, the Urban Canyons 2024 project was undertaken to study conditions at unprecedented detail over various neighborhoods in Chicago, Illinois. This dataset consists of upper air soundings that were collected as part of this effort. Soundings were launched during two intensive observing periods, IOP1 occurred on 22-23 July 2024, while IOP2 occurred on 27-28 July 2024. For IOP1, soundings were launched at coordinated times from three sites, Shedd Aquarium in Downtown Chicago, Abizu Campus High School in Humboldt Park, and Gary Comer Youth Center in West Woodlawn. For IOP2, the Gary Comer site was replaced by a neighborhood site in West Woodlawn, Chicago. The Abizu Campos site was operated by Valparaiso University and used iMET-4 rawinsondes, the other sites were operated by the University of Illinois Urbana-Champaign and used GRAW DFM-19 sondes.This dataset contains netCDF files containing quality-controlled temperature, dewpoint, geopotential height, pressure, and vector wind measurements at 1 second intervals following launch. These files are readable by the open-source netCDF software libraries available in many software packages (i.e., python, R, fortran, C++, etc.). The dataset also contains quicklook plots of each launch on a skew-T log-p thermodynamic diagram. These are in png format viewable by most web browsers.

54 ENVIRONMENTAL SCIENCES↗

Focused Ion Beam Tomography of Alloy 617 Corroded in Molten Chloride Salt

Materials qualification of reactor structural materials is a critical step in rapid implementation of advanced nuclear reactor technologies, particularly to assess the corrosion performance in these designs. Accelerated qualification of reactor structural materials requires incorporating powerful computational toolsets, such as phase field modelling in the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework, to predict the evolution of structural materials due to corrosion. Accordingly, computational toolsets will require experimental data generated at appropriate length scales to validate accuracy. Focused ion beam (FIB) provides a high degree of control over manipulation of materials for analytical purposes, including capturing data on the evolution in the microstructure and elemental composition of materials at the mesoscale, an appropriate length scale for phase field modelling of intergranular diffusion phenomena using the MOOSE framework. For instance, the FEI Helios G4 UX dual beam plasma FIB microscope at the Irradiated Materials Characterization Laboratory (IMCL) is capable of backscatter diffraction (EBSD) and energy-dispersive x-ray spectroscopy (EDS) documenting the evolution in the microstructure and elemental composition, respectively. The Helios can perform EDS and EBSD three-dimensionally (3D) using tomography, which is then combined using different software packages to visualize 3D volumes correlating elemental composition to microstructural data. The purpose of this investigation was to develop a streamlined characterization and data processing workflow for 3D tomography studies on the FEI Helios G4 plasma FIB. The investigation is segmented into three parts: 1) Optimizing the data collection workflow, 2) identifying appropriate data processing and visualization software (i.e. DREAM.3D, MIPAR, and VGStudioMax), and 3) establishing an infrastructure for public release. The optimization of the data collection workflow is in collaboration with members of the U220 department to setup formal training on the tomography operation of the G4, through ThermoFisher Scientific, and exploring DREAM.3D, MIPAR, and VGStudioMax data processing/visualization software packages. VGStudioMax currently demonstrates the most promise for future use. Optimization of the data collection and processing workflow is still ongoing. A collaboration with INL High Performance Computing (HPC) established an open-source license for expediting the public release of FIB tomography datasets through HPC. FIB tomography data generated by the G4 will provide comprehensive data for validating 3D phase field mesoscale modelling tools within the MOOSE framework for accelerated qualification of reactor structural materials.

Copeland-Johnson, Trishelle↗

MOSAIC-CONUS: A Multimodal, Multi-Temporally Paired Dataset for Earth Sciences

Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.

54 ENVIRONMENTAL SCIENCES↗

Toward the Neutrino Discovery Platform: An Auditable, Uncertainty-Bearing Toolchain for MINERvA Open-Data Cross-Section Analysis

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]↗

Observational ozone datasets over the global oceans and polar regions (version 2024)

Studying tropospheric ozone over the remote areas of the planet, such as the open oceans and the polar regions, is crucial to understand the role of ozone as a global climate forcer and regulator of atmospheric oxidative capacity. A focus on the pristine oceanic and polar regions complements the available land-based datasets and provides insights into key photochemical and depositional loss processes that control the concentrations and spatiotemporal variability in ozone as well as the physicochemical mechanisms driving these patterns. However, an assessment of the role of ozone over the oceanic and polar regions has been hampered by a lack of comprehensive observational datasets. Here, we present the first comprehensive collection of ozone data over the oceans and the polar regions. The overall dataset consists of 77 ship cruises/buoy-based observations and 48 aircraft-based campaigns. The dataset, consisting of more than 630 000 independent ozone measurement data points covering the period from 1977 to 2022 and an altitude range from the surface to 5000 m (with a focus on the lowest 2000 m), allows systematic analyses of the spatiotemporal distribution and long-term trends over the 11 defined ocean/polar regions. The datasets from ships, buoys, and aircraft are complemented by ozonesonde data from 29 launch sites or field campaigns and by 21 non-polar and 17 polar ground-based station datasets. The datasets contain information on how long the observed air masses were isolated from land, as estimated by backward trajectories from the individual observation points. To extract observations representative of oceanic conditions, we recommend using a subset of the data with an isolation time of 72 h or longer, from the analysis with coincident radon observations. These filtered oceanic and polar data showed typically flat diurnal cycles at high latitudes, whereas daytime decreases in ozone (11 %–16 %) were observed at lower latitudes. The ship/buoy- and aircraft-based datasets presented here will supplement the land-based ones in the TOAR-II (Tropospheric Ozone Assessment Report Phase II) database to provide a fully global assessment of tropospheric ozone. The described dataset is available at https://doi.org/10.17596/0004044 (Kanaya et al., 2025).

Kanaya, Yugo [Japan Agency for Marine-Earth Scienc↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

RectifHydPlus Data Pipeline

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab↗

RectifHydPlus Data Pipeline v1.1.0

The RectifHydPlus Data Pipeline is an open source and fully reproducible data processing pipeline for creating RectifHydPlus—a dataset of historical monthly net electricity generation for all US hydropower plants (>10MW). The pipeline is coded in R, applying tidyverse libraries and code principles, and using the targets data pipeline framework. All data inputs to the RectifHydPlus Data Pipeline are available from public sources. References to all data inputs, as well as instructions for running the RectifHydPlus Data Pipeline, are available on the GitLab code repository: https://code.ornl.gov/turnersw/rectifhydplus

Turner, SeanWilliam Donald [Oak Ridge National Lab↗

Assessing Spatial Representativeness of Global Flux Tower Eddy-Covariance Measurements Using Data from FLUXNET2015

Large datasets of carbon dioxide, energy, and water fluxes were measured with the eddy-covariance (EC) technique, such as FLUXNET2015. These datasets are widely used to validate remote-sensing products and benchmark models. One of the major challenges in utilizing EC-flux data is determining the spatial extent to which measurements taken at individual EC towers reflect model-grid or remote sensing pixels. To minimize the potential biases caused by the footprint-to-target area mismatch, it is important to use flux datasets with awareness of the footprint. This study analyze the spatial representativeness of global EC measurements based on the open-source FLUXNET2015 data, using the published flux footprint model (SAFE-f). The calculated annual cumulative footprint climatology (ACFC) was overlaid on land cover and vegetation index maps to create a spatial representativeness dataset of global flux towers. The dataset includes the following components: (1) the ACFC contour (ACFCC) data and areas representing 50%, 60%, 70%, and 80% ACFCC of each site, (2) the proportion of each land cover type weighted by the 80% ACFC (ACFCW), (3) the semivariogram calculated using Normalized Difference Vegetation Index (NDVI) considering the 80% ACFCW, and (4) the sensor location bias (SLB) between the 80% ACFCW and designated areas (e.g. 80% ACFCC and window sizes) proxied by NDVI. Finally, we conducted a comprehensive evaluation of the representativeness of each site from three aspects: (1) the underlying surface cover, (2) the semivariogram, and (3) the SLB between 80% ACFCW and 80% ACFCC, and categorized them into 3 levels. The goal of creating this dataset is to provide data quality guidance for international researchers to effectively utilize the FLUXNET2015 dataset in the future.

54 ENVIRONMENTAL SCIENCES↗

MINERvA s Open Data Product: A First for Neutrino Data Preservation

Access to information on neutrino nucleus interactions is critical to the success of all neutrino oscillation experiments. MINERvA's rich dataset covers a range of energies and nuclei unique amongst experiments, and as such is critical to the community in building the important shared knowledge needed to unravel the mysteries of the neutrino. In particular, its dataset provides the greatest statistical coverage in in the range of neutrino energies pertinent for DUNE until DUNE's near detector begins operation. Historically, such significant datasets in neutrino physics have been preserved primarily through their published results. While meaningful and useful, this limits the ability to explore the data to its fullest extent as new perspectives continue to form. MINERvA has undertaken a major effort to break this trend and preserve its data in a format to be as analyzable as possible from outside the collaboration. This has culminated in the officially-released MINERvA Open Data Product for the community to take advantage of and utilize. Maintaining direct access to the dataset in an analyzable form will allow new insights to continue to be extracted indefinitely. This talk will cover the contents of this product, the information included (and excluded), the tools provided to utilize the product effectively, the support MINERvA intends to provide in its use, and some lessons learned through the process.

Last, David [Rochester U.] (ORCID:0000000245147183↗

Annual and sub-seasonal dynamics of a rapidly eroding permafrost coastline along the Beaufort Sea in northern Alaska

Drew Point, an unlithified ice-rich permafrost coastline along the Alaskan Beaufort Sea, is among the most rapidly eroding Arctic coastlines, with an average erosion rate of 19 m/yr from 2007 to 2019. We use 16 high-resolution remote sensing datasets (satellite, airborne, and UAV imagery) to analyze erosion mechanisms (thermal abrasion and denudation) in relation to environmental forcings along a 1.5 km stretch of coastline during the 2018 and 2019 open water seasons. In a striking contrast, 2019 exhibited the highest mean erosion rate (34.5 m) within the 2007–2019 record, while 2018 had the second lowest (11.2 m). Block failure contributed to sub-seasonal erosion rates 6 to 21 times higher than thermal denudation, with staggered block fall timing, lag responses post-storm, and non-storm block collapse influencing overall erosion magnitude and timing. To quantify wind effects, we developed wind sums, a metric combining cumulative wind speed and directional data that can be used as a proxy for integrated storm intensity capable of incorporating lagged responses that correlated strongly with erosion at sub-seasonal and annual scales. Our findings emphasize the dominant role of wind during periods of open water and air temperature during the thaw season in driving permafrost coastline erosion dynamics, while highlighting the importance of spatiotemporally high-resolution datasets for understanding Arctic coastal change dynamics.

Alaska Beaufort Sea Coast↗

CoCoMET v1.0: a unified open-source toolkit for atmospheric object tracking and analysis

Advances in performance and analysis capabilities have accelerated the development of object tracking algorithms for atmospheric research. This has resulted in a growing number of studies using Lagrangian tracking techniques to analyze the evolution of atmospheric phenomena and the underlying processes. However, the increasing complexity and variety of tracking algorithms present a steep learning curve for new users and make it difficult for existing users to compare algorithm performance. We introduce CoCoMET (Community Cloud Model Evaluation Toolkit), an open-source toolkit that addresses these issues. CoCoMET simplifies the process of running multiple tracking algorithms simultaneously and analyzing objects in both model and observational datasets by specifying parameters in a single configuration file. It standardizes input data from different sources into a consistent format and unifies the tracking output across algorithms. CoCoMET enhances the functionality of existing tracking methods by calculating additional properties such as cell growth and dissipation rates, perimeter, surface area, convexity, and irregularity. In addition, CoCoMET includes a novel method for identifying mergers and splits in 2D and 3D tracks and supports the integration of Eulerian/stationary datasets external to the tracking data for process studies. Its potential utility is demonstrated through examples of model intercomparison, model evaluation against observations, and comparisons between tracking algorithms. Designed for open-source environments, CoCoMET will continue to expand with future releases, incorporating more input data types and tracking algorithms.

54 ENVIRONMENTAL SCIENCES↗

An analysis of parameter compression and Full-Modeling techniques with Velocileptors for DESI 2024 and beyond

In anticipation of forthcoming data releases of current and future spectroscopic surveys, we present the validation tests and analysis of systematic effects within velocileptors modeling pipeline when fitting mock data from the AbacusSummit N-body simulations. We compare the constraints obtained from parameter compression methods to the direct fitting (Full-Modeling) approaches of modeling the galaxy power spectra, and show that the ShapeFit extension to the traditional template method is consistent with the Full-Modeling method within the standard ΛCDM parameter space. We show the dependence on scale cuts when fitting the different redshift bins using the ShapeFit and Full-Modeling methods. We test the ability to jointly fit data from multiple redshift bins as well as joint analysis of the pre-reconstruction power spectrum with the post-reconstruction BAO correlation function signal. We further demonstrate the behavior of the model when opening up the parameter space beyond ΛCDM and also when combining likelihoods with external datasets, namely the Planck CMB priors. Finally, we describe different parametrization options for the galaxy bias, counterterm, and stochastic parameters, and employ the halo model in order to physically motivate suitable priors that are necessary to ensure the stability of the perturbation theory.

79 ASTRONOMY AND ASTROPHYSICS↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

IM3 Projected U.S. Western Interconnection Grid Stress Dataset

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from GO WEST and TEP) for Integrated Multisector, Multiscale Modeling (IM3) Phase 2 simulations across eight different scenarios for the U.S. Western Interconnection through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States from a set of Thermodynamic Global Warming (TGW) simulations. These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 GO WEST is an open-source power grid modeling framework for the U.S. Western Interconnection, which allows users to tailor the model depending on their research study and science questions. It covers 28 balancing authorities (BAs) and 12 states in U.S. Western Interconnection. GO WEST allows users to select different number of nodes and come up with a simplified network by utilizing 10,000 nodal topology of the U.S. Western Interconnection (ACTIVSg10k). Users can select different number of nodes, mathematical formulations (linear programming vs. mixed-integer linear programming), transmission line limit scaling factors, and hurdle rate scaling factors. GO WEST offers a unit commitment and economic dispatch (UC/ED) module to simulate grid operations on an hourly scale. In this sense, users can calibrate and validate their model versions by comparing model outputs to historical datasets. TEP is an open-source transmission capacity expansion model, built on the GO WEST framework. It utilizes linear programming to optimize transmission capacity addition investment on existing lines within the GO WEST framework. The TEP model only increases the thermal capacity of existing transmission lines and does not add new lines to the system, which leaves the topology preserved. In order to use TEP model, users need to create scenarios with the GO WEST framework. Please refer to README file for a detailed description of the dataset including individual files and references.

Capacity Expansion Model↗

Data-driven organic solubility prediction at the limit of aleatoric uncertainty

Abstract Small molecule solubility is a critically important property which affects the efficiency, environmental impact, and phase behavior of synthetic processes. Experimental determination of solubility is a time- and resource-intensive process and existing methods for in silico estimation of solubility are limited by their generality, speed, and accuracy. This work presents two models derived from the FASTPROP and CHEMPROP architectures and trained on BigSolDB which are capable of predicting solubility at arbitrary temperatures for a wide range of small molecules in organic solvent. Both extrapolate to unseen solutes 2–3 times more accurately than the current state-of-the-art model and we demonstrate that they are approaching the aleatoric limit (0.5–1$$\log S$$ log S ) of available test data, suggesting that further improvements in prediction accuracy require more accurate datasets. The FASTPROP-derived model (called FASTSOLV) and the CHEMPROP-based model are open source, freely accessible via a Python package and web interface, highly reproducible, and up to 2 orders of magnitude faster than current alternatives.

Science & Technology - Other Topics↗