A Framework to Predict Variability Characteristics in Building Load Profiles
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Having the ability to predict the protein-encoding gene content of an incomplete genome or metagenome-assembled genome is important for a variety of bioinformatic tasks. In this study, as a proof of concept, we built machine learning classifiers for predicting variable gene content in Escherichia coli genomes using only the nucleotide k-mers from a set of 100 conserved genes as features. Protein families were used to define orthologs, and a single classifier was built for predicting the presence or absence of each protein family occurring in 10%–90% of all E. coli genomes. The resulting set of 3,259 extreme gradient boosting classifiers had a per-genome average macro F1 score of 0.944 [0.943–0.945, 95% CI]. We show that the F1 scores are stable across multi-locus sequence types and that the trend can be recapitulated by sampling a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including “hypothetical proteins” was accurately predicted (F1 = 0.902 [0.898–0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions had slightly lower F1 scores but were still accurate (F1s = 0.895, 0.872, 0.824, and 0.841 for transposon, phage, plasmid, and antimicrobial resistance-related functions, respectively). Finally, using a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources, we observed an average per-genome F1 score of 0.880 [0.876–0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data.
Abstract Predicting forced, long‐term radiative feedbacks from internal climate variability has been a decades‐long quest in climate science. We train a convolutional neural network (CNN) to predict annual‐ and global‐mean top of the atmosphere radiation anomalies from time‐varying maps of near‐surface temperature in climate models. Trained on internal variability alone, the nonlinear CNN can predict radiation under strong climate change, outperforms a regularized linear regression approach, and works within and across different climate models. We show with explainable artificial intelligence methods that the CNN draws predictive skill from physically meaningful regions but at much smaller spatial scales than currently assumed.
Explore the source record for details and available documents.
We present a probabilistic framework tailored for solar energy applications referred to as the Weather Research and Forecasting-Solar ensemble prediction system (WRF-Solar EPS). WRF-Solar EPS has been developed by introducing stochastic perturbations into the most relevant physical variables for solar irradiance predictions. In this study, we comprehensively discuss the impact of the stochastic perturbations of WRF-Solar EPS on solar irradiance forecasting compared to a deterministic WRF-Solar prediction (WRF-Solar DET), a stochastic ensemble using the stochastic kinetic energy backscatter scheme (SKEBS), and a WRF-Solar multi-physics ensemble (WRF-Solar PHYS). The performances of the four forecasts are evaluated using irradiance retrievals from the National Solar Radiation Database (NSRDB) over the contiguous United States. We focus on the predictability of the day-ahead solar irradiance forecasts during the year of 2018. The results show that the ensemble forecasts improve the quality of the forecasts, compared to the deterministic prediction system, by accounting for the uncertainty derived by the ensemble members. However, the three ensemble systems are under-dispersive, producing unreliable and overconfident forecasts due to a lack of calibration. In particular, WRF-Solar EPS produces less optically thick clouds than the other forecasts, which explains the larger positive bias in WRF-Solar EPS (31.7 W/m 2 ) than in the other models (22.7–23.6 W/m 2 ). This study confirms that the WRF-Solar EPS reduced the forecast error by 7.5% in terms of the mean absolute error (MAE) compared to WRF-Solar DET, and provides in-depth comparisons of forecast abilities with the conventional scientific probabilistic approaches (i.e., SKEBS and a multi-physics ensemble). Guidelines for improving the performance of WRF-Solar EPS in the future are provided.
Deep Learning (DL) models are increasingly used throughout the sciences. However, their performance and usefulness depend greatly on their architecture which is defined by hyperparameters such as the number of nodes, layers, the learning rate, etc. Tuning these hyperparameters is time-consuming because evaluating their performance requires a lengthy training step. Stochastic optimizers used in training lead to performance variability and potentially prediction reliability issues. In this talk, we will describe an automated optimization method based on surrogate models and active learning strategies for tuning DL model architectures. We take into account the prediction variability with the goal to identify architectures that make reliable and robust predictions. We demonstrate our developments on an application arising in particle physics.
Rainfed agriculture is the mainstay of economies across Southern Africa (SA), where most precipitation is received during the austral summer monsoon. This study aims to further our understanding of monsoon precipitation predictability over SA. We use three natural climate forcings, El Niño–Southern Oscillation, Indian Ocean Dipole (IOD), and the Indian Ocean Precipitation Dipole (IOPD)—the dominant precipitation variability mode—to construct an empirical model that exhibits significant skill over SA during monsoon in explaining precipitation variability and in forecasting it with a five-month lead. While most explained precipitation variance (50%–75%) comes from contemporaneous IOD and IOPD, preconditioning all three forcings is key in predicting monsoon precipitation with a zero to five-month lead. Seasonal forecasting systems accurately represent the interplay of the three forcings but show varying skills in representing their teleconnection over SA. This makes them less effective at predicting monsoon precipitation than the empirical model.
Machine-learned interatomic potentials (MLIPs) have become the state-of-the-art for performing accurate, scalable molecular dynamics (MD) simulations. It is therefore crucial to understand and quantify the reliability of MLIPs for downstream property predictions. Uncertainty in predicted properties can arise from limitations in first-principles training data, intrinsic MLIP model errors in representing the data, and the statistical noise introduced during subsequent MD simulations. Using ion transport in Li7P3S11 as a case study, we systematically assess the impact of training set size and selection, neural network stochasticity, and MD sampling statistics on predicted diffusivity and activation energy. We find that when using equivariant MLIP architectures with standard MD protocols, uncertainty arising from MD sampling dominates over model-induced errors. In contrast, MLIP errors relative to the underlying first-principles data are consistently minor. Given this, there are two main routes to improving the accuracy of predictions based on MLIP potentials: adopting higher accuracy reference data generation methods, and improving the MD sampling statistics.
This article is a Commentary on Robbins et al . (2024), 244 : 2239–2250 .
The El Niño-Southern Oscillation (ENSO) influences climate variability globally, encompassing various other modes of variability, and thus represents a key predictable climate signal on seasonal timescales. Yet, its response to greenhouse warming remains uncertain, with models projecting a range of outcomes. Here, we demonstrate that in response to warming, a state-of-the-art high-resolution climate model simulates a rapid transition from a moderate-amplitude irregular regime, as observed in the current climate, to a highly regular oscillation with intensifying amplitude. This behaviour can be attributed to increasing air-sea feedbacks, which approach criticality in the second half of this century, and growing atmospheric noise. As ENSO intensifies in this model, it synchronizes with other prominent climate modes, such as the North Atlantic Oscillation and the Indian Ocean Dipole, thereby imprinting its regular, predictable variability on them. If realized, this global climate mode resonance would have wide-ranging whiplash impacts on regional hydroclimates.
Abstract Distributions of both native and invasive species are expected to shift under future climate. Species distribution models (SDMs) are often used to explore future habitats, but sources of uncertainty including novel climate conditions may reduce the reliability of future projections. We explore the potential spread of the invasive annual grass ventenata ( Ventenata dubia ) in the western United States under both current and future climate scenarios using boosted regression tree models and 30 global climate models (GCMs). We quantify novel climate conditions, prediction variability arising from both the SDMs and GCMs, and the agreement among GCMs. Results demonstrate that currently suitable habitat is concentrated inside the invaded range of the northwest, but substantial habitat exists outside the invaded range in the Southern Rockies and southwestern US mountains. Future suitability projections vary greatly among GCMs, but GCMs commonly projected decreased suitability in the invaded range and increased suitability along higher elevations of interior mountainous areas. Climate novelty did not appear to undermine the prediction reliability in many cases where the climate–species relationship was fully represented by the occurrence data. GCM‐derived variability resulting from variation in future cool season precipitation and temperature seasonality was greatest in the Rocky Mountains. SDM‐derived variability was higher in currently suitable habitat, and few GCMs projections agreed that these areas would contain future suitable habitat. However, while prediction variability was high, many GCM projections agreed that parts of the Rocky, Wasatch, and Uinta Mountains would contain highly suitable habitat in the future. As disturbances in the interior mountains occur in coming decades, reducing some natural barriers to invasion, land managers, and conservationists will need to monitor for ventenata in post‐disturbance environments. Changes to invasion potential may not play out for several decades, but results related to current potential may have applications for early detection and rapid response planning.
With the increased use of data-driven approaches and machine learning-based methods in material science, the importance of reliable uncertainty quantification (UQ) of the predicted variables for informed decision-making cannot be overstated. UQ in material property prediction poses unique challenges, including multi-scale and multi-physics nature of materials, intricate interactions between numerous factors, limited availability of large curated datasets, etc. In this work, we introduce a physics-informed Bayesian Neural Networks (BNNs) approach for UQ, which integrates knowledge from governing laws in materials to guide the models toward physically consistent predictions. To evaluate the approach, we present case studies for predicting the creep rupture life of steel alloys. Experimental validation with three datasets of creep tests demonstrates that this method produces point predictions and uncertainty estimations that are competitive or exceed the performance of conventional UQ methods such as Gaussian Process Regression. Additionally, we evaluate the suitability of employing UQ in an active learning scenario and report competitive performance. The most promising framework for creep life prediction is BNNs based on Markov Chain Monte Carlo approximation of the posterior distribution of network parameters, as it provided more reliable results in comparison to BNNs based on variational inference approximation or related NNs with probabilistic outputs.
The release of leachates from intact coal ash impoundments is a concern due to the enrichment and mobilization of toxic elements such as arsenic (As) and selenium (Se). This study aims to explore the intrinsic properties of coal fly ash that correlate with the relative leachability of As and Se. We performed leaching experiments with 52 fly ash samples collected from 15 different U.S. power plants and representing coal feedstocks from the three major domestic regions. We assessed the mobilization potential of As and Se in fly ash based on standardized leaching protocols and performed multivariate and lasso regression analyses to explore correlations of leachable As and Se contents with characteristics such as major element contents, loss on ignition, and pH. The results of regression models indicated that major elements (Fe, Ca, and Al) for a wide range of fly ashes can serve as predictor variables for the leaching potential of As but not for Se. LOI and pH were not important predictive variables in the models. Both regression approaches resulted in relatively strong fits for leachable As (correlation coefficient R 2 = 0.78 for both models) compared to models for leachable Se (R 2 = 0.49). Overall, these results suggest that correlation models combined with on-site elemental analysis with portable analyzers may enable a screening method for leachable As in coal ash.
Electromagnetic ion cyclotron (EMIC) waves play a key role in radiation belt dynamics through resonant interactions. However, their low occurrence probability, high variability, and spatial intermittency pose challenges for accurate modeling. In this study, we present a machine learning (ML)-based global EMIC wave model built on the entire data set from the Van Allen Probes mission. To capture the distinct statistical characteristics of wave occurrence and amplitude, the model is separated into two modules: an occurrence model trained using ML techniques, and a wave amplitude model sampled from observed probability distributions. The input parameters are limited to real-time or predictable variables to ensure practical applicability. Our model shows strong performance across the entire test set and demonstrates improved predictive capability over a baseline random occurrence model, particularly during quiet geomagnetic conditions. Evaluation during both quiet and active periods confirms the model's ability to represent the clustered and intermittent nature of EMIC wave activity. Furthermore, the model provides global estimates of wave power, enabling integration with radiation belt electron data and showing signatures consistent with wave-induced scattering. We found a good correlation between the global wave activity from the model and relativistic electron observation by Van Allen Probes, regardless of the availability of in situ wave observations. The modular structure of the model also allows for straightforward expansion for additional wave properties, such as wave frequency, which can be modeled independently. This flexible, event-sensitive approach offers a promising framework for data-driven radiation belt simulations and space weather applications.
Abstract. Large-scale interaction between the three tropical ocean basins is an area of intense research that is often conducted through experimentation with numerical models. A common problem is that modeling groups use different experimental setups, which makes it difficult to compare results and delineate the role of model biases from differences in experimental setups. To address this issue, an experimental protocol for examining interaction between the tropical basins is introduced. The Tropical Basin Interaction Model Intercomparison Project (TBIMIP) consists of experiments in which sea surface temperatures (SSTs) are prescribed to follow observed values in selected basins. There are two types of experiments. One type, called standard pacemaker, consists of simulations in which SSTs are restored to observations in selected basins during a historical simulation. The other type, called pacemaker hindcast, consists of seasonal hindcast simulations in which SSTs are restored to observations during 12-month forecast periods. TBIMIP is coordinated by the Climate and Ocean – Variability, Predictability, and Change (CLIVAR) Research Focus on Tropical Basin Interaction. The datasets from the model simulations will be made available to the community to facilitate and stimulate research on tropical basin interaction and its role in seasonal-to-decadal variability and climate change.
Flavodoxins (Flds) mediate the flux of electrons between oxidoreductases in diverse metabolic pathways. To investigate whether Flds can support electron transfer to a sulfite reductase (SIR) that evolved to couple with a ferredoxin, we evaluated the ability of Flds to transfer electrons from a ferredoxin-NADP reductase (FNR) to a ferredoxin-dependent SIR using growth complementation of an Escherichia coli strain with a sulfur metabolism defect. We show that Flds from cyanobacteria complement this growth defect when coexpressed with an FNR and an SIR that evolved to couple with a plant ferredoxin. When we evaluated the effect of peptide insertion on Fld-mediated electron transfer, we observed a sensitivity to insertions within regions predicted to be proximal to the cofactor and partner binding sites, while a high insertion tolerance was detected within loops distal from the cofactor and within regions of helices and sheets that are proximal to those loops. Bioinformatic analysis showed that natural Fld sequence variability predicts a large fraction of the motifs that tolerate insertion of the octapeptide SGRPGSLS. In conclusion, these results represent the first evidence that Flds can support electron transfer to assimilatory SIRs, and they suggest that the pattern of insertion tolerance is influenced by interactions with oxidoreductase partners.
CATALYST proposes to perform foundational coordinated research in a team-oriented collaborative effort aimed at advancing a robust understanding of modes of Earth system variability and change using models, observations and process studies. The proposed research will address the DOE/BER mission by exploring the limits to predictability, identifying fundamental underlying mechanisms, quantifying interactions among modes of variability, and discovering tipping points in the Earth system to understand the current and future impacts of these phenomena on regional and global climate. Four fundamental gaps are identified in our knowledge of the Earth system: 1) What are the limits to predictability on various timescales? 2) What are the interactions among modes of Earth system variability? 3) How may modes of Earth system variability change in response to changes in external forcing, and what are the tipping points involved with those changes? 4) How are high impact events connected to modes of Earth system variability and how may they change in the future? Related to those gaps in our knowledge, we formulate four research objectives to address those gaps using a combination of Earth system models (ESMs) and machine learning (ML) methods. Research Objective 1 (RO1) addresses the first gap above and proposes to understand modes of variability and their limits of predictability on subseasonal to decadal timescales using ESMs and ML. Research Objective 2 (RO2) addresses the second gap and proposes to use a hierarchy of models to understand relevant processes and feedbacks related to how modes of variability interact with each other. Research Objective 3 (RO3) is designed to study the third gap and proposes to examine the role of external forcings in changes of modes of Earth system variability and their interactions, and the likelihood and predictability of tipping points and irreversible changes. Research Objective 4 (RO4) will address the fourth gap and proposes to use high resolution ESMs, regionally refined models (RRMs), and ML methods to investigate the relationships between high impact events (e.g. flash droughts and precipitation extremes, atmospheric rivers (ARs), tropical cyclones (TCs), storm surge/sea level rise), the synoptic systems that produce them, and their changes related to modes of Earth system variability. The research will involve the use of the Community Earth System Model (CESM), Energy Exascale Earth System Model (E3SM), CMIP multi-model data sets, a hierarchy of simpler models, and numerous observational data sets. In the course of the proposed research, CATALYST will contribute to metrics and diagnostics that will be integrated in Coordinated Model Evaluation Capabilities (CMEC), particularly with regards to the Quasi-biennial Oscillation (QBO) and its interactions with the Madden-Julian Oscillation (MJO), high atmospheric pressure blocking, and new precipitation metrics.
Model predictive control (MPC) has been widely studied as a promising approach for improving energy efficiency and operational flexibility in buildings, yet its real-world performance for commercial variable air volume (VAV) systems remains insufficiently characterized. In particular, the impacts of model mismatch on control robustness, real-time computational burden, and device-level operation are rarely evaluated using long-term field data. Here, this study presents a comprehensive experimental evaluation of MPC applied to a full-scale VAV system in Oak Ridge National Laboratory’s Flexible Research Platform-2 building with constant cooling/heating temperature setpoints and no occupancy. The study offers three key advantages over existing work: (1) it uses a representative building in a full-scale experimental test, capturing realistic system dynamics and complexity; (2) it evaluates a relatively sophisticated MPC formulation using two different optimization solvers (Gurobi and PSO), fully accounting for computational complexity and methodological diversity; and (3) it systematically assesses potential negative impacts on various building devices, benchmark against a well-established baseline, ASHRAE Guideline 36 (G36). To isolate zone- and air-handling-unit–level supervisory control effects, the supply fan was operated with a fixed static pressure setpoint under all strategies, and the trim-and-response static pressure reset in G36 was not enabled. Results show that MPC maintained thermal comfort while improving energy efficiency. Abrupt solar radiation variations degraded performance. Computation times ranged from ∼1 s (Gurobi) to ∼ 70 s (PSO). Compared with G36, MPC achieves 33% energy savings and reduces median reheat coil output by approximately a factor of 5–10 for a representative cooling day under matched weather conditions. However, it increases the maximum discomfort deviation from 0.5 to 1°C and results in a 32% increase in staging frequency. In addition, PSO-based MPC introduced damper oscillations, also affecting actuator longevity.