Leveraging Active Subspaces to Capture Epistemic Model Uncertainty in Deep Generative Models for Molecular Design
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Numerical weather prediction (NWP) models, such as the Weather Research and Forecasting (WRF) model, are widely used to provide estimates of the offshore wind energy resource owing to their large spatial coverage compared to available observations. Nevertheless, spatiotemporal distribution of model biases is highly dependent on factors including model configuration, location, and the interplay of multi-scale physical processes. Here, in this study, we focus on the characterization of model uncertainties in simulated coast-to-offshore winds over the northeast U.S., by varying sea surface temperature (SST) forcings, surface layer (SL) and planetary boundary layer (PBL) parameterizations, as well as identifying biases that may be directly passed from initial and boundary conditions. Multiple measurements, including aircraft data collected during the U.S. Department of Energy's Two-Column Aerosol Project (TCAP) experiment, are used to constrain the model results and facilitate quantitative comparisons. Our analysis indicates while SST forcing has notable impacts on simulated air temperature and moisture within PBL, the modeled winds are in general more sensitive to the choices of SL and PBL physics than to SST. The model’s forcing data not only controls the vertical dependence of wind speed errors, but also alters regional variability in wind speed’s spatial correlation. Bias comparisons between ERA5 reanalysis and ensemble simulations revealed significant similarity, particularly in wind speed biases during winter, underscoring their dependency on initial and boundary conditions. Coastal and offshore near-surface wind speed biases tend to exhibit much higher similarity in winter than in summer due to the presence of much stronger and more persistent synoptic wind conditions. This study highlights the importance of accurate atmospheric forcing and parameterization choices in improving wind forecasts and suggests the potential for extrapolating coastal wind biases to offshore locations, aiding wind energy forecasting and informing the Wind Forecast Improvement Project-3 (WFIP3).
We present a simple comparative framework for testing and developing uncertainty modeling in uncertain marching cubes implementations. The selection of a model to represent the probability distribution of uncertain values directly influences the memory use, run time, and accuracy of an uncertainty visualization algorithm. We use an entropy calculation directly on ensemble data to establish an expected result and then compare the entropy from various probability models, including uniform, Gaussian, histogram, and quantile models. Our results verify that models matching the distribution of the ensemble indeed match the entropy. We further show that fewer bins in nonparametric histogram models are more effective whereas large numbers of bins in quantile models approach data accuracy.
This dataset contains model output and input data, as well as source code examples for the Terrestrial Ecosystem Model with the Dynamic Vegetation Model and Dynamic Organic Soil (DVM-DOS-TEM) for the field sites Imnavait creek and the Bonanza creek Long Term Ecological Research Network (LTER). The data covers simulations from the last glacial maximum (LGM) until 2100 for a selection of paleo scenarios, setting the mean temperature of the LGM up to 10°C lower than pre-industrial conditions. The model structure was modulated to represent various model versions, and this dataset contains the relevant changes in the source code. The raw output data, the processed statistical data, the setup and processing scripts as well as parameter value distribution files from a parameter sensitivity analysis are included as well. Model outputs include active layer depth, organic soil carbon, soil layer depths, gross primary productivity (GPP) with and without nitrogen limitation, net primary productivity (NPP), soil liquid water content, heterotrophic, maintenance, and growth respiration, soil temperature, and vegetation carbon (*.nc files). The Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) project is a research effort to reduce uncertainty in the Department of Energy’s Energy Exascale Earth System Model (E3SM) by developing a predictive understanding of Arctic tundra ecosystems underlain by permafrost and to quantify feedbacks from the Arctic tundra to the Earth system. NGEE Arctic is supported by the Department of Energy's Office of Biological and Environmental Research.Over Phases 1–3, observations made by the NGEE Arctic team across a gradient of permafrost landscapes in Arctic Alaska improved the representation of tundra processes in the land surface component of E3SM (the E3SM Land Model, ELM). Model improvements emphasized unique aspects of permafrost environments and explored reductions in model complexity while retaining predictive power. The Arctic-informed ELM developed by NGEE Arctic has been used to make novel predictions on processes ranging from permafrost thaw to soil biogeochemical cycling to Earth system feedbacks associated with the unique characteristics of tundra plants. In Phase 4, the NGEE Arctic team is evaluating our new predictive understanding under novel conditions across the Arctic domain. In collaboration with partners at long-term pan-Arctic research sites we are examining whether an Arctic-informed ELM can faithfully simulate interactions among surface and subsurface processes at site, regional, and pan-Arctic scales. In turn, we are using variety of tools to dynamically extend and evaluate ELM inference, with an emphasis on data synthesis and pan-Arctic model evaluation, reintegration of code with an evolving E3SM, scaling across heterogeneous Arctic landscapes, and the appropriate representation of the impacts of increasingly frequent Arctic disturbances.
The precise atomic structure and therefore the wavelength-dependent opacities of lanthanides are highly uncertain. This uncertainty introduces systematic errors in modeling transients like kilonovae and estimating key properties such as mass, characteristic velocity, and heavy metal content. Here, we quantify how atomic data from across the literature as well as choices of thermalization efficiency of r-process radioactive decay heating impact the light curve and spectra of kilonovae. Specifically, we analyze the spectra of a grid of models produced by the radiative transfer code Sedona that span the expected range of kilonova properties to identify regions with the highest systematic uncertainty. Our findings indicate that differences in atomic data have a substantial impact on estimates of lanthanide mass fraction, spanning approximately 1 order of magnitude for lanthanide-rich ejecta, and demonstrate the difficulty in precisely measuring the lanthanide fraction in lanthanide-poor ejecta. Mass estimates vary typically by 25%–40% for differing atomic data. Similarly, the choice of thermalization efficiency can affect mass estimates by 20%–50%. Observational properties such as color and decay rate are highly model dependent. Velocity estimation, when fitting solely based on the light curve, can have a typical error of ∼100%. Atomic data of light r-process elements can strongly affect blue emission. Even for well-observed events like GW170817, the total lanthanide production estimated using different atomic data sets can vary by a factor of ∼6.
Aerosols mediate the radiative fluxes in clear and cloudy skies and dominate the uncertainty in the radiative forcing of climate. Relatively dense networks of aerosol optical depth measurements have been used to effectively constrain simulated aerosol optical properties, but model diversity of aerosol microphysical properties is much larger (Mann et al. 2014, Myhre et al. 2009). A better understanding of aerosol microphysics is essential as they are the fundamental pieces of information required to convert aerosol emissions information to radiative and cloud-nucleating properties that drive radiative effects and forcing of climate. Model representation of aerosol microphysical properties has lagged in part due to their inherently greater spatial variability, but also due to a lack of available observations. This project – the Printed Optical Particle Spectrometer Network (POPSnet)-Southern Great Plains (SGP) Pilot – launched the first spatially dense network of aerosol size distribution measurements over an area the size of a global model grid cell and demonstrated its use in providing valuable information regarding the spatial variability of aerosol microphysical properties and its drivers for improving model constraints.
Not Available
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
This study examines the modeling uncertainty of wind resource data stemming from the use of various planetary boundary layer (PBL) parameterizations available in the Weather Research and Forecasting (WRF) model. WRF-based wind simulations spanning 20 years at 3-km resolution using 11 different PBL schemes are used to objectively investigate the uncertainty in modeling wind speed for land-based wind (LBW) and offshore wind (OSW) locations in Puerto Rico. The uncertainty in the wind modeling for the 20-year dataset is quantified using the spread index (SI) and standard deviation (SD). For virtual LBW and OSW sites, the SI and SD values are analyzed as calculated across various spatial and temporal scales. Because the PBL's atmospheric stability conditions can be characterized into two dominant categories, the study focuses on analyzing the SI and SD for daytime (mainly unstable PBL conditions) and nighttime (mainly stable PBL conditions). For wind shear (10 m-200 m) at the OSW and LBW sites, WRF-based numerical experiments indicate the following SI (or SD) ranges: 39%-94% (0.74 m/s-1.44 m/s) during the daytime for OSW, 50%-75% (0.68 m/s-1.19 m/s) during the daytime for LBW, 37%-60% (0.73 m/s-1.12 m/s) during the nighttime for OSW, and 57%-143 % (0.65 m/s-1.43 m/s) during the nighttime for LBW. While a high SI is observed when modeling LBW during the nighttime, there are notable modeling uncertainties during the daytime on the leeward side of the orographic barriers for Puerto Rico.
Earth system models (ESMs) and general circulation models (GCMs) are heavily used to provide inputs to sectoral impact and multisector dynamic models, which include representations of energy, water, land, economics, and their interactions. Therefore, representing the full range of model uncertainty, scenario uncertainty, and interannual variability that ensembles of these models capture is critical to the exploration of the future co-evolution of the integrated human–Earth system. The pre-eminent source of these ensembles has been the Coupled Model Intercomparison Project (CMIP). With more modeling centers participating in each new CMIP phase, the size of the model archive is rapidly increasing, which can be intractable for impact modelers to effectively utilize due to computational constraints and the challenges of analyzing large datasets. In this work, we present a method to select a subset of the latest phase, CMIP6, featuring models for use as inputs to a sectoral impact or multisector dynamics models, while prioritizing preservation of the range of model uncertainty, scenario uncertainty, and interannual variability in the full CMIP6 ensemble results. This method is intended to help impact modelers select climate information from the CMIP archive efficiently for use in downstream models that require global coverage of climate information. This is particularly critical for large-ensemble experiments of multisector dynamic models that may be varying additional features beyond climate inputs in a factorial design, thus putting constraints on the number of climate simulations that can be used. We focus on temperature and precipitation outputs of CMIP6 models, as these are two of the most used variables among impact models, and many other key input variables for impacts are at least correlated with one or both of temperature and precipitation (e.g., relative humidity). Besides preserving the multi-model ensemble variance characteristics, we prioritize selecting CMIP6 models in the subset that preserve the very likely distribution of equilibrium climate sensitivity values as assessed by the latest Intergovernmental Panel on Climate Change (IPCC) report. This approach could be applied to other output variables of climate models and, possibly when combined with emulators, offers a flexible framework for designing more efficient experiments on human-relevant climate impacts. It can also provide greater insight into the properties of existing CMIP6 models.
Uncertainty visualization is a key component in translating important insights from ensemble simulation data into actionable decision-making by visually conveying various aspects of uncertainty within a system. With the recent advent of fast surrogate models trained on ensemble data, we can substitute computationally expensive simulations, which allows users to interact with more aspects of data spaces than ever before. However, the use of ensemble data with surrogate models in a decision-making tool brings up new challenges for uncertainty visualization, namely how to reconcile and communicate the new and different types of uncertainties brought in by surrogates and how to utilize these new data estimates in actionable ways. In this work, we examine these issues as they relate to high-dimensional data visualization, the integration of discrete datasets and the continuous representations of those datasets, and the unique difficulties associated with systems that allow users to iterate between input and output spaces. We assess the role of uncertainty visualization in facilitating intuitive and actionable interaction with ensemble data and surrogate models, and highlight key challenges in this new frontier of computational simulation.
Surrogate models are a critical ingredient to computation-based design and validation of many DOE mission-relevant physical systems. When first-principles computation of properties of a physical systems becomes pro hibitive, surrogate models are the only path towards achieving tasks such as uncertainty quantification (UQ), exploration of design space, and validation of design choices. In this project we have developed and demonstrated a new surro gate modeling paradigm for complex models that is data-driven, non-intrusive, and has the potential to be versatile and equipped with performance guaran tees. This combination of features is absent in existing surrogate modeling tools. The framework we have developed in this project exploits a quantum-classical correspondence to establish a quantum system that mimics the dynamics of the classical Hamiltonian system from which data in the form of temporal snapshots is provided. Since quantum dynamics propagates distributions over observables, the framework is naturally suited to propagation of epistemic uncertainties in the form of distributions over initial state and parametric uncertainties. In this project, we take the first step in establishing this novel framework by deriving a quantization and de-quantization procedure, demonstrating the accuracy of the quantum surrogate models these define using two model systems, and defining the next steps in maturing the framework towards a tool applicable to Sandia mission-relevant problems.
Tropical Cyclones (TCs) are intense storms that pose a persistent and considerable risk to coastal communities and infrastructure in the global tropics and subtropics, including the United States (US). With known limitations associated with observations and high-resolution earth system models, synthetic TC models that capture a wide spectrum of storm possibilities have been developed to robustly quantify TC risk. Here we examine the simulation of various TC features in the North Atlantic relevant for US coastal risk in three synthetic TC models forced with ERA5 reanalysis: MIT, CHAZ and RAFT. While there is a broad agreement among these models in terms of their representation of salient TC characteristics, certain differences do exist. To connect these modeling uncertainties with energy infrastructure resilience, we apply fragility curves that link simulated TC intensities to damage probabilities, demonstrating how uncertainty in storm states may translate into that in coastal impacts. Our study indicates that acknowledging and accounting for inter-model uncertainty leads to more reliable risk assessments, strengthening science-to-action pathways for managing risks associated with TCs.
Uncertainties in the thermophysical properties of molten salts impact both the steady-state and transient behavior of Molten Salt Reactors (MSRs). In this work, we aim to quantify the influence of such uncertainties on the transient operation of the Molten Chloride Reactor Experiment (MCRE), utilizing the open-source specifications provided for this reactor. Seven representative transient scenarios are considered. For each scenario, we evaluate the impact of thermophysical property uncertainties on four key multiphysics model output variables of interest (VoIs): maximum power density, maximum fuel temperature, maximum reflector temperature, and average fuel velocity magnitude. In addition, we perform a Global Sensitivity Analysis (GSA) by computing Sobol’ indices for the uncertain input parameters to determine their contribution to the variability of each VoI. Conducting GSA is computationally intensive due to the large number of required evaluations of the high-fidelity multiphysics model. To mitigate this cost, we develop a surrogate modeling framework that combines Gaussian Process (GP) regression with Principal Component Analysis (PCA), enabling efficient sample generation for the GSA. Our results show that for energy-related VoIs, thermal conductivity is the dominant contributor to uncertainty. In contrast, for flow-related VoIs, density and dynamic viscosity are the primary sources of uncertainty. The specific heat of the fuel salt was found to play a secondary role in the transient analyses.
Traffic emissions significantly impact near-road air quality and public health. This research applies a Bayesian modeling framework to investigate these impacts using high-resolution traffic and air pollutant data from an urban corridor in Columbia, South Carolina. Despite a data collection period truncated by the COVID-19 lockdown, the Bayesian approach successfully identified significant predictors and quantified model uncertainty. Employing Bayesian Model Selection and Averaging enhanced prediction accuracy and evaluated model uncertainty. Findings indicate that higher temperatures and increased moisture levels elevate particulate matter (PM 1.0 , PM 2.5 , PM 10 ) concentrations, while traffic speed significantly affects nitrogen dioxide (NO 2 ) levels. Specifically, higher average traffic speeds (indicative of smoother flow) correspond to lower NO 2 concentrations, suggesting that less congested conditions reduce NO 2 emissions. This study highlights the robustness of Bayesian methods for generating reliable air quality insights even under data-constrained conditions. The findings underscore the importance of traffic flow management (e.g., reducing congestion) for mitigating near-road NO 2 exposure and provide a basis for developing targeted public health strategies.
This report describes challenges associated with the hierarchical Bayesian approach to inform model-form uncertainty (MFU) representations, which are parameterized modifications to a mathematical models’ governing equations to express uncertainty in form of the equations. To inform model-form uncertainties, hierarchical Bayesian inference is often employed. Here, the MFU parameters are distributed parametrically, and the hyperparameters of the parametric distribution are informed through Bayesian inference, with the aim of determining the MFU parameter distribution that best agrees with calibration data. In practice, however, we have found the hierarchical Bayesian approach falls short of this aim. We discuss theoretical and methodological challenges of the approach, and we present several numerical demonstrations of these challenges. To conclude, we suggest promising alternative approaches for future investigation.