Search NASA⌕ Search

SEARCH · Search NASA

Results for “ensemble”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Sensitivity Analysis of Drivers Water Shortage in the Los Angeles Region During Drought

The code and detailed step-by-step instructions for generating the model output data, processing results, and analysis and plotting are provided at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. The PyArtes model is a python adaptation of the Artes model. PyArtes uses many of the same input data and optimization model architecture as Artes. Documentation for the PyArtes model is provided in the Supplement to the paper. The primary data product are simulated monthly water shortages for indoor and outdoor demand under a large ensemble of drought scenarios (>13,000). The droughts are hypothetical and are not based on historical time series data of supply sources - though historical data did help inform ranges explored for supply parameters. Demands are informed by recent 2017-2021 water supply data. Demands used for the model can be accessed at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. Simulations resolve demand for over 90 water providers in the study region. The results report 36 months of water shortage data for each indoor and outdoor demand node. The study also developed a multilayer perceptron (MLP) neural network trained on a subset of the simulated shortage ensemble to emulate worst annual water shortage for a given set of parameter multipliers -- provided the parameter values fall within the ranges sampled in the ensemble. Emulated water shortages for synthetic ensembles are in the MLP-generated shortages folder. The MLP model was used to generate larger ensembles to support Sobol analysis that would have been extremely computationally expensive to simulate. Datasets provided in this repository*: Simulated shortages. These results are used for the analysis for Figures 5, 8, and 9 in the paper, and also to train the MLP emulator. .zip file containing outputs for the 13,312 scenario ensemble. Separate .csv files for indoor and outdoor shortage for each scenario. Rows = demand ids (~100), Columns = months (36) Units = acre-feet/month of shortage (shortage = monthly demand - supply). 1 acft = 1233.48 m^3 .csv files of aggregated shortages derived from the 13,312 ensemble Rows = scenarios (13,312), Columns = demand ids (~100) Units = acre-feet/year (either worst annual shortage or total shortage over the 3-year drought) .csv file of the parameter multipliers scenarios for the ensemble .csv file of the parameter ranges and baseline values the multipliers were applied to MLP-generated shortages. These results are used for Figures 4, 6, and 7 in the paper. mwd higher folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results Emulated shortages. Rows = scenarios, columns = demand ids, units acft Sobol results. Rows = demand ids, columns Sobol (S1, ST, or 95% confidence interval) value for each parameter mwd lower folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results same organization as mwd higher MLP performance: performance metrics (R^2, RMSE, BIAS, MAPE) for the testing subset (20% or 2,662 scenarios) and simulated vs emulated worst year shortage (acre-feet/year) for every demand node, MWD wholesale regions, and the entire study region (LAC). Supporting data for figures. Figure plotting scripts in the associated GitHub repo. These files support analysis and visualization. Geospatial Data used for plotting simulated water shortages and Sobol results. Dictionary of full names for demand nodes in the model and estimates of water supply by source type informed by Artes input files and California Urban Water Management Planning data: https://water.ca.gov/Programs/Water-Use-And-Efficiency/Urban-Water-Use-Efficiency/Urban-Water-Management-Plans *Readme files provided for each folder.

drought↗

Estimating Uncertainty in Simulated ENSO Statistics

Abstract Large ensembles of model simulations are frequently used to reduce the impact of internal variability when evaluating climate models and assessing climate change induced trends. However, the optimal number of ensemble members required to distinguish model biases and climate change signals from internal variability varies across models and metrics. Here we analyze the mean, variance and skewness of precipitation and sea surface temperature in the eastern equatorial Pacific region often used to describe the El Niño–Southern Oscillation (ENSO), obtained from large ensembles of Coupled model intercomparison project phase 6 climate simulations. Leveraging established statistical theory, we develop and assess equations to estimate, a priori, the ensemble size or simulation length required to limit sampling‐based uncertainties in ENSO statistics to within a desired tolerance. Our results confirm that the uncertainty of these statistics decreases with the square root of the time series length and/or ensemble size. Moreover, we demonstrate that uncertainties of these statistics are generally comparable when computed using either pre‐industrial control or historical runs. This suggests that pre‐industrial runs can sometimes be used to estimate the expected uncertainty of statistics computed from an existing historical member or ensemble, and the number of simulation years (run duration and/or ensemble size) required to adequately characterize the statistic. This advance allows us to use existing simulations (e.g., control runs that are performed during model development) to design ensembles that can sufficiently limit diagnostic uncertainties arising from simulated internal variability. These results may well be applicable to variables and regions beyond ENSO.

54 ENVIRONMENTAL SCIENCES↗

Approximate CFTs and random tensor models

Abstract A key issue in both the field of quantum chaos and quantum gravity is an effective description of chaotic conformal field theories (CFTs), that is CFTs that have a quantum ergodic limit. We develop a framework incorporating the constraints of conformal symmetry and locality, allowing the definition of ensembles of ‘CFT data’. These ensembles take on the same role as the ensembles of random Hamiltonians in more conventional quantum ergodic phases of many-body quantum systems. To describe individual members of the ensembles, we introduce the notion of approximate CFT, defined as a collection of ‘CFT data’ satisfying the usual CFT constraints approximately, i.e. up to small deviations. We show that they generically exist by providing concrete examples. Ensembles of approximate CFTs are very natural in holography, as every member of the ensemble is indistinguishable from a true CFT for low-energy probes that only have access to information from semi-classical gravity. To specify these ensembles, we impose successively higher moments of the CFT constraints. Lastly, we propose a theory of pure gravity in AdS 3 as a random matrix/tensor model implementing approximate CFT constraints. This tensor model is the maximum ignorance ensemble compatible with conformal symmetry, crossing invariance, and a primary gap to the black-hole threshold. The resulting theory is a random matrix/tensor model governed by the Virasoro 6j-symbol.

Physics↗

Identifying Northern Hemisphere Stratospheric and Surface Temperature Responses to the Mt. Pinatubo Eruption within E3SMv2-SPA

The Mt. Pinatubo eruption on 15 June 1991 is often associated with surface warming in the subsequent Northern Hemisphere winter. Employing E3SMv2 with prognostic aerosol modifications, we generated an ensemble of simulations initialized on 1 June 1991 to limit the intra-ensemble variability at the time of the eruption and a more traditional ensemble representing the full range of intra-ensemble variability. For each ensemble member we generated a paired counterfactual simulation with the Pinatub forcing removed allowing for isolation of the Pinatubo impact. In general, the limited variability ensemble has greater coherence in the Pinatubo impact across ensemble members which leads to more statistically robust signals compared to the full variability ensemble. Stratospheric warming patterns from Pinatubo were approximately zonally symmetric and confined between 30°S and 50°N. Isolating localized surface temperature impacts was more difficult, but the limited variability simulation did identify a preferential region of cooling between 20°S to 50°N.

54 ENVIRONMENTAL SCIENCES↗

A Practical Probabilistic Benchmark for AI Weather Models

Since the weather is chaotic, it is necessary to forecast an ensemble of future states. Recently, multiple AI weather models have emerged claiming breakthroughs in deterministic skill. Unfortunately, it is hard to fairly compare ensembles of AI forecasts because variations in ensembling methodology become confounding and the baseline data volume is immense. We address this by scoring lagged initial condition ensembles—whereby an ensemble can be constructed from a library of deterministic hindcasts. This allows the first parameter‐free intercomparison of leading AI weather models' probabilistic skill against an operational baseline. Lagged ensembles of the two leading AI weather models, GraphCast and Pangu, perform similarly even though the former outperforms the latter in deterministic scoring. These results are elaborated upon by sensitivity tests showing that commonly used multiple time‐step loss functions damage ensemble calibration.

54 ENVIRONMENTAL SCIENCES↗

Improving seasonal precipitation forecasts in the Western United States through statistical downscaling

Abstract Seasonal precipitation forecasts in the western United States are critical resources for water resource management, especially during winter. While current seasonal forecasting systems provide monthly precipitation forecasts operationally, their coarse resolution limits their effectiveness in capturing the localized precipitation patterns and snowpack conditions essential for water resource managers in the mountainous regions. Here, analog statistical downscaling is demonstrated as an effective approach to enhance the spatial resolution of operational seasonal forecasts provided by the North American Multi-Model Ensemble. Downscaling was performed by building an analog ‘library’, in which corresponding model forecasts and observed values during the training period were stored. In the testing period, unseen model forecasts referenced the closest historical forecast from the analog library and applied the corresponding observational value for each point. This analysis indicates that downscaled products can capture localized features more accurately than the original coarse resolution forecasts, reducing forecast error across the western United States. Moreover, downscaling individual ensemble members—rather than downscaling the ensemble mean—further reduces forecasting error for their multi-model ensemble mean products. The greatest error reductions in the downscaled product, measured by root mean squared error (RMSE), were observed at low to mid-elevations (500–2000 meters), with 50%–70% improvement relative to the original forecasts. In the higher elevations (2000 meters and above), changes in RMSE relative to the original forecast were limited to 10%–30% improvements. The improvement is more substantial for forecast systems with 10 ensemble members compared to that with 4 members, but this relationship does not hold for the system with 24 ensemble members. These findings show that analog statistical downscaling can effectively address the spatial limitations of seasonal precipitation forecasts with minimal computational cost, providing a valuable framework for enhancing coarse resolution forecasting products while providing insights into the timing of ensemble mean calculations during the downscaling process.

Vernon, B. (ORCID:0009000891670689)↗

CONUS-wide Projected Flood Frequency and Uncertainty Estimates, Version 1.0

This dataset presents a large-ensemble of CONUS-wide projected flood frequency and uncertainty estimates across ~2.7 million NHDPlusV2 river reaches over the CONUS. The framework producing this dataset leverages a multi-model, uncertainty-aware modeling framework that allows evaluating shifts in flood frequences at the stream reach level across the CONUS. CONUS-wide ensemble streamflow projections generated from hydrologic simulations driven by downscaled and bias-corrected Coupled Model Intercomparison Project Phase 6 (CMIP6) outputs are used to derive these flood frequency and uncertainty estimates over the period 1980 - 2099. A spatially consistent regional L-moment algorithm is applied across clusters defined by the US Hydrologic Unit Code Subregions (HUC4s and HUC8s) and NHDPlusV2 stream orders to estimate flood frequencies. The dataset also includes at-site based flood estimates that allow for the comparison between local and regional approach-based estimates, assess projected changes, and characterize their uncertainties. For more reliable estimation of rare flood frequencies such as 500 and 1000-year return periods, super-ensemble based estimates are also included in the dataset. This dataset is derived to support the "Impact-Informed Dam Safety Risk Assessment for Securing Hydropower Assests" project for the US Department of Energy (DOE) Hydropower and Hydrokinetic Office (H2O). For further details, refer to Kao et al. (2022), Ghimire et al. (2023), Ghimire et al. (2025), and Hosking and Wallis (1997).

Ghimire, Ganesh [ORNL] (ORCID:0000000242843941)↗

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Evaluating the Importance of Conformers for Understanding the Vacuum-Ultraviolet Spectra of Oxiranes: Experiment and Theory

Vacuum-ultraviolet (VUV) absorption spectroscopy enables electronic transitions that offer the unambiguous identification of molecules. As target molecules become more complex, multifunctional species present a great challenge to both experimental and computational spectroscopy. This research reports both experimental and theoretical studies of oxiranes. Computationally, the nuclear ensemble approach has been used to accurately predict experimental spectra for a variety of molecules. However, this approach incurs great computational cost, as ensembles generally consist of thousands of geometries. The present study aims to drastically reduce the ensemble by evaluating the significance of the conformers to the predicted spectra. This approach was applied to 11 substituted oxiranes using the Conformer Rotamer Ensemble Sampling Tool (CREST) of Grimme to generate an ensemble of unique conformers determined by their Boltzmann populations. Five TD-DFT functionals (BMK, CAM-B3LYP, M06-2X, MN15, ωB97X-D) and EOM-CCSD were used to simulate the spectrum of each substituted oxirane ensemble. Computed spectra were then compared to the experiment using both qualitative and quantitative metrics. Based on these metrics, it was observed that certain conformers may not be necessary to characterize this set of oxiranes despite the temperature (323 K) of the experiment. A single conformer can then be used with TD-DFT and EOM-CCSD to replicate the experimental spectra of these medium-sized combustion species.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

RASPA3

RASPA3, a molecular simulation code for computing adsorption and diffusion in nanoporous materials and thermodynamic and transport properties of fluids. It implements force field based classical Monte Carlo/molecular dynamics in various ensembles. RASPA3 is rewritten from the ground up in C++23 with speed and code readability in mind. Transition-matrix Monte Carlo is added to compute the density of states and free energies. The Monte Carlo code for rigid molecules is based on quaternions, and the atomic positions needed in the energy evaluation are recreated from the center of mass position and quaternion orientation. The expanded ensemble methodology for fractional molecules, with a scaling parameter λ between 0 and 1, now also keeps track of analytic expressions of dU/dλ, allowing independent verification of the chemical potential using thermodynamic integration. The source code is freely available under the MIT license on GitHub.

Dubbeldam, David↗

Maximum Entropy Principle in Deep Thermalization and in Hilbert-Space Ergodicity

We report universal statistical properties displayed by ensembles of pure states that naturally emerge in quantum many-body systems. Specifically, two classes of state ensembles are considered: those formed by (i) the temporal trajectory of a quantum state under unitary evolution or (ii) the quantum states of small subsystems obtained by partial, local projective measurements performed on their complements. These cases, respectively, exemplify the phenomena of “Hilbert-space ergodicity” and “deep thermalization.” In both cases, the resultant ensembles are defined by a simple principle: The distributions of pure states have maximum entropy, subject to constraints such as energy conservation, and effective constraints imposed by thermalization. We present and numerically verify quantifiable signatures of this principle by deriving explicit formulas for all statistical moments of the ensembles, proving the necessary and sufficient conditions for such universality under widely accepted assumptions, and describing their measurable consequences in experiments. We further discuss information-theoretic implications of the universality: Our ensembles have maximal information content while being maximally difficult to interrogate, establishing that generic quantum state ensembles that occur in nature hide (scramble) information as strongly as possible. Our results generalize the notions of Hilbert-space ergodicity to time-independent Hamiltonian dynamics and deep thermalization from infinite to finite effective temperature. Our work presents new perspectives to characterize and understand universal behaviors of quantum dynamics using statistical and information-theoretic tools.

Eigenstate thermalization↗

Data and Scripts Associated with "Modeling Ecohydrological Responses of Vegetation to Urban Microclimates Using the E3SM Land Model"

This dataset supports the study of vegetation ecohydrological responses to urban microclimates using the land component of the Energy Exascale Earth System Model (ELM) at four urban sites in Knoxville, Tennessee, USA. It includes the model inputs, simulation outputs, and associated scripts for running ELM simulations and analyzing the resulting data. The Model_Inputs folder includes static surface data, satellite-derived phenology (i.e., leaf area index), and atmospheric forcing data used to drive ELM simulations. Detailed descriptions of these datasets are provided in Section 2.3.2 of the associated manuscript. The Model_Outputs folder contains simulation results for the baseline, treatment, and ensemble experiments. Outputs from the baseline and treatment simulations are provided as raw ELM NetCDF files. Because the raw outputs from the 4,000-member ensemble are prohibitively large, the ensemble results are provided as summarized CSV files, which also serve as the source data for Figure 5 of the associated manuscript. The Scripts folder contains three components: E3SM, the core codebase of the Energy Exascale Earth System Model (E3SM); elm-olmt, the Offline Land Model Testbed (OLMT) used to perform the simulations; and knoxville_elm, which contains the analysis scripts used to process model outputs and generate the figures and results presented in the associated manuscript. Additional information is provided in Scripts_readme.txt within the Scripts directory.

Lu, Xiaoman [ORNL] (ORCID:0000000306698780)↗

Enhancing Fluid Flow Pressure and Saturation Prediction Accuracy and Reducing Uncertainty with Committee Machine – Illinois Basin Decatur Project (IBDP) as a Case Study

Presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. Carbon capture and storage (CCS) is a way to play a critical role in the global transition to a low-emission economy. Current progress is hampered by a number of factors, among which the lack of risk-informed design tools and decision support frameworks is seen as a major roadblock. Significant interest exists in using artificial intelligence to accelerate CCS site feasibility studies, as well as to facilitate the permit application process. Existing works commonly train a single deep learning model. This work investigates the feasibility of using a conventional ensemble learning (committee machine) technique to further improve prediction accuracy. Ensemble-based algorithms generally improve over individual base learners in terms of robustness and accuracy. Deep ensembles, however, are time-consuming to create and train. A pragmatic question is whether small-sized ensembles may lead to prediction improvement. Here we evaluated the efficacy of an ensemble learning technique using the latent spectral model (LSM), an efficient deep neural operator algorithm, as base learners. Preliminary results, obtained using the Illinois Basin-Decatur Project (IBDP) carbon sequestration data/model, show that small-sized ensembles can improve prediction over the base learners, achieving prediction accuracy of ~1.6 psi root mean square error (RMSE) on pressure (relative the average reservoir pressure of 3150 psi), and less than 1.3% for saturation.

Sun, Alexander↗

Arctic Impact Identification with Less Data Using Variable Relationships: An Exploratory Express LDRD project.

Regional impacts from sea ice loss can be challenging to separate from internal climate variability, potentially requiring thousands of ensemble members. East Asian wintertime cooling has been linked to sea ice loss from present day conditions in the Polar Amplification Model Intercomparison Project with these large ensemble counts. This cooling is theorized to arise from a strengthened Siberian High and East Asian Jet response. The strengthened Siberian High can be detected with one fifth the ensemble members needed for the East Asian wintertime cooling in a single model. We thus hypothesize that leveraging relationships between multiple variables in a conditional pathways-based approach would reduce the number of required ensemble members to conclusively attribute East Asian wintertime cooling to future sea ice concentrations. In all analyzed cases, confidence was increased when evaluating sea ice loss’s responsibility for the joint effects of East Asian cooling, East Asian Jet strengthening, and Siberian High strengthening over just East Asian cooling. However, we were not able to confidently attribute future East Asian wintertime cooling to sea ice loss in a single model. We found that significant intra-ensemble variability within single Earth System Models (ESMs) produced highly uncertain forcing response models upon which attribution results were undermined. We were able to show that ensemble mean seasonally averaged metrics from multiple ESMs greatly improved the accuracy of the forcing response linear models and exposed the necessity of all three steps in the pathway (sea ice area, Siberian High pressure, and East Asian Jet speed) for accurate prediction of East Asian wintertime cooling. Although all three steps were necessary, East Asian wintertime cooling possesses a large dependence on the Siberian High pressure, which weakens the confidence associated with overall strong joint-attribution comparing present day and future scenarios. We believe transitioning the pathway nodes to relative changes between the Siberian High and Aleutian Low as well as between the midlatitude westerlies and subtropical jet in the East Asianj Jet region may be able to produce significant attribution more fully dependent upon all three steps. Ultimately, this research demonstrates the simple extensibility of conditional pathways-based attribution to sea ice loss forcing on the Earth system.

54 ENVIRONMENTAL SCIENCES↗

How well are hazards associated with derechos reproduced in regional climate simulations?

Abstract. A 15-member ensemble of convection-permitting regional simulations of the fast-moving and destructive derecho of 29–30 June 2012 that impacted the northeastern urban corridor of the USA is presented. This event generated 1100 reports of damaging winds, generated significant wind gusts over an extensive area of up to 500 000 km2, caused several fatalities, and resulted in widespread loss of electrical power. Extreme events such as this are increasingly being used within pseudo-global-warming experiments to examine the sensitivity of historical, societally important events to global climate non-stationarity and how they may evolve as a result of changing thermodynamic and dynamic contexts. As such it is important to examine the fidelity with which such events are described in hindcast experiments. The regional simulations presented herein are performed using the Weather Research and Forecasting (WRF) model. The resulting ensemble is used to explore simulation fidelity relative to observations for wind gust magnitudes, spatial scales of convection (as is manifest in high composite reflectivity, cREF), and both rainfall and hail production as a function of model configuration (microphysics parameterization, lateral boundary conditions (LBCs), start date, use of nudging, compiler choice, damping, and number of vertical levels). We also examine the degree to which each ensemble member differs with respect to key mesoscale drivers of convective systems (e.g., convective available potential energy and vertical wind shear) and critical manifestations of deep convection, e.g., vertical velocities, cold-pool generation, and how those properties relate to the correct characterization of the associated atmospheric hazards (wind gusts and hail). Use of a double-moment, seven-class scheme with number concentrations for all species (including hail and graupel) results in the greatest fidelity of model-simulated wind gusts and convective structure to the observations of this event. All ensemble members, however, fail to capture the intensity of the event in terms of the spatial extent of convection and the production of high near-surface wind gusts. We further show very high sensitivity to the LBCs employed and specifically that simulation fidelity is higher for simulations nested within ERA-Interim compared to ERA5. Excess convective available potential energy (CAPE) in all ensemble members after the derecho passage leads to excess production of convective cells, wind gusts, cREF > 40 dBZ, and precipitation during a frontal passage on the subsequent day. This event proved very challenging to forecast in real time and to reproduce in the 15-member hindcast simulation ensemble presented here. Future work could examine if simulations with other initial and lateral boundary conditions can achieve greater fidelity.

Shepherd, Tristan (ORCID:0000000186276419)↗

Enhancing Fluid Flow Pressure and Saturation Prediction Accuracy and Reducing Uncertainty with Committee Machine – Illinois Basin Decatur Project (IBDP) as a Case Study

This is the conference paper accompanying an oral presentation at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24, 2024. Carbon capture and storage (CCS) is a way to play a critical role in the global transition to a low-emission economy. Current progress is hampered by a number of factors, among which the lack of risk-informed design tools and decision support frameworks is seen as a major roadblock. Significant interest exists in using artificial intelligence to accelerate CCS site feasibility studies, as well as to facilitate the permit application process. Existing works commonly train a single deep learning model. This work investigates the feasibility of using a conventional ensemble learning (committee machine) technique to further improve prediction accuracy. Ensemble-based algorithms generally improve over individual base learners in terms of robustness and accuracy. Deep ensembles, however, are time-consuming to create and train. A pragmatic question is whether small-sized ensembles may lead to prediction improvement. Here we evaluated the efficacy of an ensemble learning technique using the latent spectral model (LSM), an efficient deep neural operator algorithm, as base learners. Preliminary results, obtained using the Illinois Basin-Decatur Project (IBDP) carbon sequestration data/model, show that small-sized ensembles can improve prediction over the base learners, achieving prediction accuracy of ~1.6 psi root mean square error (RMSE) on pressure (relative the average reservoir pressure of 3150 psi), and less than 1.3% for saturation.

Sun, Alexander↗

Evaluation of the 2022 West Nile virus forecasting challenge, USA

Abstract Background West Nile virus (WNV) is the most common cause of mosquito-borne disease in the continental USA, with an average of ~1200 severe, neuroinvasive cases reported annually from 2005 to 2021 (range 386–2873). Despite this burden, efforts to forecast WNV disease to inform public health measures to reduce disease incidence have had limited success. Here, we analyze forecasts submitted to the 2022 WNV Forecasting Challenge, a follow-up to the 2020 WNV Forecasting Challenge. Methods Forecasting teams submitted probabilistic forecasts of annual West Nile virus neuroinvasive disease (WNND) cases for each county in the continental USA for the 2022 WNV season. We assessed the skill of team-specific forecasts, baseline forecasts, and an ensemble created from team-specific forecasts. We then characterized the impact of model characteristics and county-specific contextual factors (e.g., population) on forecast skill. Results Ensemble forecasts for 2022 anticipated a season at or below median long-term WNND incidence for nearly all (> 99%) counties. More counties reported higher case numbers than anticipated by the ensemble forecast median, but national caseload (826) was well below the 10-year median (1386). Forecast skill was highest for the ensemble forecast, though the historical negative binomial baseline model and several team-submitted forecasts had similar forecast skill. Forecasts utilizing regression-based frameworks tended to have more skill than those that did not and models using climate, mosquito surveillance, demographic, or avian data had less skill than those that did not, potentially due to overfitting. County-contextual analysis showed strong relationships with the number of years that WNND had been reported and permutation entropy (historical variability). Evaluations based on weighted interval score and logarithmic scoring metrics produced similar results. Conclusions The relative success of the ensemble forecast, the best forecast for 2022, suggests potential gains in community ability to forecast WNV, an improvement from the 2020 Challenge. Similar to the previous challenge, however, our results indicate that skill was still limited with general underprediction despite a relative low incidence year. Potential opportunities for improvement include refining mechanistic approaches, integrating additional data sources, and considering different approaches for areas with and without previous cases. Graphical Abstract

54 ENVIRONMENTAL SCIENCES↗

Solid-state platform for cooperative quantum dynamics driven by correlated emission

While traditionally regarded as an obstacle to quantum coherence, recent breakthroughs in quantum optics have shown that the dissipative interaction of a qubit with its environment can be leveraged to protect quantum states and synthesize many-body entanglement. Inspired by this progress, here we set the stage for the—as yet uncharted—exploration of analogous cooperative phenomena in hybrid solid-state platforms. We develop a comprehensive formalism for the quantum many-body dynamics of an ensemble of solid-state spin defects interacting with the magnetic field fluctuations of a common solid-state reservoir. Our framework applies to any solid-state reservoir whose fluctuating spin, pseudospin, or charge degrees of freedom generate magnetic fields. To understand whether correlations induced by dissipative processes can play a relevant role in a realistic experimental setup, we apply our model to a qubit array interacting via the spin fluctuations of a ferromagnetic bath. Our results show that the low-temperature collective relaxation rates of the qubit ensemble can display clear signatures of super- and subradiance, i.e., forms of cooperative dynamics traditionally achieved in atomic ensembles. We find that the solid-state analog of these cooperative phenomena is robust against spatial disorder in the qubit ensemble and thermal fluctuations of the magnetic reservoir, providing a route for their feasibility in near-term experiments. Finally, our work lays the foundation for a multiqubit approach to quantum sensing of solid-state systems and the direct generation of many-body entanglement in spin-defect ensembles. Furthermore, we discuss how the tunability of solid-state reservoirs opens up novel pathways for exploring cooperative phenomena in regimes beyond the reach of conventional quantum optics setups.

NV centers↗