Search NASA⌕ Search

SEARCH · Search NASA

Results for “synthetic forecast”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Error-Level-Controlled Synthetic Forecasts for Renewable Generation

Renewable energy resources, including solar and wind energy, play a significant role in sustainable energy systems. However, the inherent uncertainty and intermittency of renewable generation pose challenges to the safe and efficient operation of power systems. Recognizing the importance of short-term (hours ahead) renewable generation forecasting in power systems operation, it becomes crucial to address the potential inaccuracies in these forecasts. To systematically evaluate the performance of controllers in the presence of imperfect forecasts, we generate synthetic forecasts using actual renewable generation profiles (one from solar and one from wind). These synthetic forecasts incorporate different levels of statistical error, allowing us to control and manipulate the accuracy of the predictions. The primary objective is to employ synthetic forecasts with controlled yet realistic error levels to systematically investigate how controllers adapt to variations in forecast accuracy, providing valuable insights into their robustness and effectiveness under real-world conditions.

Array↗

Quantifying and simulating the weather forecast uncertainty for advanced building control

Weather forecast uncertainty is unavoidable despite technological advancements. Accurately quantifying and modelling this uncertainty is essential for developing and comparing advanced building controllers. In this study, we present a structured approach using a first-order autoregressive model (AR(1)) to model uncertainty in ambient temperature and global solar irradiation (GHI) forecasts. We analyzed weather data from four cities and employed Jensen–Shannon divergence (JSD) to evaluate the similarity between synthetic and actual forecast errors. The average JSD values for temperature are 0.027 (Berkeley), 0.021 (Leuven), 0.018 (Berlin), and 0.008 (Oslo), and for GHI, the average JSD values are 0.016 (Berkeley), 0.058 (Leuven), and 0.013 (Berlin). The low JSD values indicate a high similarity between the synthetic and real forecast error distributions. Further, our approach successfully generates synthetic weather forecasts that mirror the statistical properties of actual forecasts. The implementation of our method for uncertain forecast generation is being added to the BOPTEST framework.

54 ENVIRONMENTAL SCIENCES↗

Time-series forecasting using manifold learning, radial basis function interpolation, and geometric harmonics

We address a three-tier numerical framework based on nonlinear manifold learning for the forecasting of high-dimensional time series, relaxing the “curse of dimensionality” related to the training phase of surrogate/machine learning models. At the first step, we embed the high-dimensional time series into a reduced low-dimensional space using nonlinear manifold learning (local linear embedding and parsimonious diffusion maps). Then, we construct reduced-order surrogate models on the manifold (here, for our illustrations, we used multivariate autoregressive and Gaussian process regression models) to forecast the embedded dynamics. Finally, we solve the pre-image problem, thus lifting the embedded time series back to the original high-dimensional space using radial basis function interpolation and geometric harmonics. The proposed numerical data-driven scheme can also be applied as a reduced-order model procedure for the numerical solution/propagation of the (transient) dynamics of partial differential equations (PDEs). In conclusion, we assess the performance of the proposed scheme via three different families of problems: (a) the forecasting of synthetic time series generated by three simplistic linear and weakly nonlinear stochastic models resembling electroencephalography signals, (b) the prediction/propagation of the solution profiles of a linear parabolic PDE and the Brusselator model (a set of two nonlinear parabolic PDEs), and (c) the forecasting of a real-world data set containing daily time series of ten key foreign exchange rates spanning the time period 3 September 2001–29 October 2020.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The Potential Benefits of Handling Mixture Statistics via a Bi-Gaussian EnKF: Tests With All-Sky Satellite Infrared Radiances

The meteorological characteristics of cloudy atmospheric columns can be very different from their clear counterparts. Thus, when a forecast ensemble is uncertain about the presence/absence of clouds at a specific atmospheric column (i.e., some members are clear while others are cloudy), that column's ensemble statistics will contain a mixture of clear and cloudy statistics. Such mixtures are inconsistent with the ensemble data assimilation algorithms currently used in numerical weather prediction. Hence, ensemble data assimilation algorithms that can handle such mixtures can potentially outperform currently used algorithms. In this study, we demonstrate the potential benefits of addressing such mixtures through a bi-Gaussian extension of the ensemble Kalman filter (BGEnKF). The BGEnKF is compared against the commonly used ensemble Kalman filter (EnKF) using perfect model observing system simulated experiments (OSSEs) with a realistic weather model (the Weather Research and Forecast model). Synthetic all-sky infrared radiance observations are assimilated in this study. In these OSSEs, the BGEnKF outperforms the EnKF in terms of the horizontal wind components, temperature, specific humidity, and simulated upper tropospheric water vapor channel infrared brightness temperatures. This study is one of the first to demonstrate the potential of a Gaussian mixture model EnKF with a realistic weather model. Our results thus motivate future research toward improving numerical Earth system predictions though explicitly handling mixture statistics.

54 ENVIRONMENTAL SCIENCES↗

Synthetic method of analogues for emerging infectious disease forecasting

The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.

97 MATHEMATICS AND COMPUTING↗

A Machine Learning Bias Correction on Large–Scale Environment of High–Impact Weather Systems in E3SM Atmosphere Model

Large–scale dynamical and thermodynamical processes are common environmental drivers of high–impact weather systems causing extreme weather events. However, such large–scale environmental conditions often display systematic biases in climate simulations, posing challenges to evaluating high–impact weather systems and extreme weather events. In this paper, a machine learning (ML) approach was employed to bias correct the large–scale wind, temperature, and humidity simulated by the atmospheric component of the Energy Exascale Earth System Model (E3SM) at ~1° resolution. The usefulness of the ML approach for extreme weather analysis was demonstrated with a focus on three high–impact weather systems, including tropical cyclones (TCs), extratropical cyclones (ETCs), and atmospheric rivers (ARs). We show that the ML model can effectively reduce climate bias in large–scale wind, temperature, and humidity while preserving their responses to imposed climate change perturbations. The bias correction is found to directly improve water vapor transport associated with ARs, and representations of thermodynamical flows associated with ETCs. When the bias–corrected large–scale winds are used to drive a synthetic TC track forecast model over the Atlantic basin, the resulting TC track density agrees better with that of the TC track model driven by observed winds. In addition, the ML model insignificantly interferes with the mean climate change signals of large–scale storm environments as well as the occurrence and intensity of three weather systems. This study suggests that the proposed ML approach can be used to improve the downscaling of extreme weather events by providing more realistic large–scale storm environments simulated by low–resolution climate models.

54 ENVIRONMENTAL SCIENCES↗

Synthetic spectra for Lyman- α forest analysis in the Dark Energy Spectroscopic Instrument

Synthetic data sets are used in cosmology to test analysis procedures, to verify that systematic errors are well understood and to demonstrate that measurements are unbiased. In this work we describe the methods used to generate synthetic datasets of Lyman-α quasar spectra aimed for studies with the Dark Energy Spectroscopic Instrument (DESI). In particular, we focus on demonstrating that our simulations reproduces important features of real samples, making them suitable to test the analysis methods to be used in DESI and to place limits on systematic effects on measurements of Baryon Acoustic Oscillations (BAO). We present a set of mocks that reproduce the statistical properties of the DESI early data set with good agreement. Additionally, we use a synthetic dataset to forecast the BAO scale constraining power of the completed DESI survey through the Lyman-α forest.

79 ASTRONOMY AND ASTROPHYSICS↗

A machine learning approach to water quality forecasts and sensor network expansion: Case study in the Wabash River Basin, United States

Abstract Midwestern cities require forecasts of surface nitrate loads to bring additional treatment processes online or activate alternative water supplies. Concurrently, networks of nitrate monitoring stations are being deployed in river basins, co‐locating water quality observations with established stream gauges. However, tools to evaluate the future value of expanded networks to improve water quality forecasts remains challenging. Here, we construct a synthetic data set of stream discharge and nitrate for the Wabash River Basin—one of the United States’ most nutrient polluted basins—using the established Agro‐IBIS and THMB models. Synthetic data enables rapid, unbiased and low‐cost assessment of potential sensor placements to support management objectives, such as near‐term forecasting. Using the synthetic data, we established baseline 1‐day forecasts for surface water nitrate at 12 cities in the basin using support vector machine regression (SVMR; RMSE 0.48–3.3 ppm). Next, we used the SVMRs to evaluate the improvement in forecast performance associated with deployment of additional nitrate sensors. We identified the optimal sensor placement to improve forecasts at each city, and the relative value of sensors at each candidate location. Finally, we assessed the co‐benefit realized by other cities when a sensor is deployed to optimize a forecast at one city, finding significant positive externalities in all cases. Ultimately, our study explores the potential for machine learning to make near‐term predictions and critically evaluate the improvement realized by expanding a monitoring network. While we use nitrate pollution in the Wabash River Basin as a case study, this approach could be readily applied to any problem where the future value of sensors and network design are being evaluated.

54 ENVIRONMENTAL SCIENCES↗

Propagating synthetic populations with dynamic Bayesian networks: a framework for long-horizon demographic forecasting

This study presents a dynamic demographic microsimulator using dynamic Bayesian networks to forecast long–term changes in household and individual life events. Leveraging longitudinal Panel Study of Income Dynamics (PSID) data, two networks for individuals and households were modeled to simulate transitions in employment, income, education, marriage, childbirth, leaving the parental home, home ownership, mortality, and household formation or dissolution. Across 1,000 simulation runs spanning 24 years, household–level outcomes remain highly accurate and individual–level predictions reasonable. Although accuracy naturally declines with projection horizon, performance remains promising at both levels. This study addresses a key limitation of existing population synthesis models, which typically generate only a single static snapshot of the population. In conclusion, by introducing a framework that propagates cross-sectional outputs into the future, the microsimulator enables the tracking of demographic evolution over time, enhances realism in population-based simulations, and supplies credible inputs to agent-based travel demand models.

Demographic modeling↗

Photovoltaic Analysis and Response Support (PARS) Platform for Solar Situational Awareness and Resiliency Services

The project's primary objective is to develop a digital-twin based Photovoltaic (PV) Analysis and Response Support (PARS) platform, which aims to provide real-time situational awareness and optimal response plans. This platform is designed to enhance the performance of hybrid PV systems, making them competitive with or even superior to conventional generation resources. The PARS platform enabled the project team to develop and evaluate an extensive suite of grid support functionalities for the hybrid PV systems to enhance grid performance, across key areas including visibility, dispatchability, security, resilience, and reliability. Given the global push toward achieving 100% clean energy by 2035, there is a significant increase in the integration of inverter-based resources (IBRs) throughout the energy grid. Effectively managing the inherent variability and uncertainty associated with IBRs is crucial for ensuring cost-effectiveness, reliability, and security in both the main grid and islanded microgrids. Constrained to a limited array of IEEE test systems or standard feeder models, traditional IBR modeling struggles to assimilate new field data, accurately reflect system dynamics, and adapt to the evolving energy landscape. In our project, we embraced a Digital Twin (DT) strategy for crafting the PARS platform. A digital twin acts as a precise virtual counterpart of a physical system, built on historical data and continuously honed with real-time insights. This enables the high-fidelity DT to accurately mirror current system operations and forecast future scenarios. Consequently, the PARS platform becomes an ideal environment for testing and refining monitoring, control, power, and energy management algorithms designed to boost hybrid PV system performance. The defining feature of the PARS platform, distinguishing it from other advanced simulation tools, is its exceptional adaptability. This is achieved by employing actual network topologies and utilizing real-time field data for fine-tuning and calibration, ensuring a close emulation of real-world conditions. The project deliverables include: 1) High-fidelity IBR models and tools for real-time parameterization, utilizing real-time field measurements to refine IBR models for enhanced accuracy and performance; 2) Grid-forming and Grid-following capabilities to deliver resilience services, including blackstart, voltage and frequency support, cold-load pick-up, power reserves, and three-phase load balancing across grid-connected and microgrid settings; 3) Machine learning-based forecasting tools and methods for generating synthetic data and topologies, creating diverse and realistic simulation environments for evaluating varied operational scenarios; 4) Advanced microgrid power and energy management algorithms for optimizing the integration and operation of PV, storage, and demand response resources within both feeder and community scales. The power grid data sets are provided by four utility companies in North Carolina and the New York Power Administration. Acting as industry advisors, our industry partners communicated stakeholder needs and regulatory standards to the research teams, aiding technology transfer by incorporating the developed methodologies into their daily operations. This collaboration ensures that the PARS platform, functioning as a power system digital twin, enhances our understanding of IBR dynamic behaviors and enables the development and evaluation of IBR control functions that match or exceed the capabilities of conventional synchronous generators.

14 SOLAR ENERGY↗

Data efficiency assessment of generative adversarial networks in energy applications

This study investigates the data requirements of generative artificial intelligence (AI), particularly generative adversarial networks (GANs), for reliable data augmentation in energy applications. Generative AI, though seen as a solution to data limitations, requires substantial data to learn meaningful distributions—a challenge often overlooked. This study addresses the challenge through synthetic data generation for critical heat flux (CHF) and power grid demand, focusing on renewable and nuclear energy. Two variants of GAN employed are conditional GAN (cGAN) and Wasserstein GAN (wGAN). Our findings include the strong dependency of GAN on data size, with performance declining on smaller datasets and varying performance when generalizing to unseen experiments. Mass flux and heated length significantly influence CHF predictions. wGAN is more robust to feature exclusion, making it suitable for constrained synthetic data generation. In energy demand forecasting, wGAN performed well for solar, wind, and load predictions. Longer lookback hours and larger datasets improved predictions, especially for load power. Seasonal variations posed challenges, with wGAN achieving a relatively high error of Root Mean Squared Error (RMSE) of 0.32 for load power prediction, compared to RMSE of 0.07 under same-season conditions. Feature exclusions impacted cGAN the most, while wGAN showed greater robustness. This study concludes that, while generative AI is effective for data augmentation, it requires substantial data and careful training to generate realistic synthetic data and generalize to new experiments in engineering applications.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

The Lyman- α forest from LBGs: First 3D correlation measurement with DESI and prospects for cosmology

The Lyman-α (Lyα) forest is a key tracer of large-scale structure at redshifts z > 2, traditionally studied using the spectra of luminous but relatively rare quasars. In this work, we explore the viability of using the fainter yet significantly more abundant Lyman Break Galaxies (LBGs) as alternative background sources for Lyα forest studies. We analyze 4,151 Lyα forest skewers extracted from LBG spectra obtained in the DESI pilot surveys conducted in the COSMOS and XMM-LSS fields. From this dataset, we present the first measurement of the Lyα forest auto-correlation function derived exclusively from LBG spectra, probing comoving separations up to 48 h -1 Mpc at an effective redshift of z eff = 2.70. The measured LBG Lyα forest auto-correlation is consistent with that derived from DESI DR2 quasar Lyα forest spectra at a comparable redshift, validating the use of LBGs as reliable background sources for Lyα forest analyses. In addition, we measure the cross-correlation between the LBG Lyα forest and the positions of 13,362 galaxies, demonstrating that this observable serves as a sensitive diagnostic for assessing the precision and accuracy of galaxy redshift estimates, and for identifying and correcting systematic offsets. Finally, using both synthetic LBG spectra and Fisher matrix forecasts, we show that a future wide-area survey covering ∼5,000 deg 2 , targeting 1,000 LBGs per square degree at signal-to-noise levels comparable to our sample, could enable LBG-based Lyα forest baryon acoustic oscillation (BAO) measurements with expected uncertainties of σ α ISO = 0.4% (isotropic) and σ α AP = 1.3% (Alcock-Paczynski). This performance is further enhanced when combining the BAO analysis with a Lyα forest Full Shape (FS) approach, yielding a predicted uncertainty of σ α ISO FS = 0.6%. These results open a new avenue for precision cosmology at high redshift using the Lyα forest in dense LBG samples.

Lyman alpha forest↗

GOOML Big Kahuna Forecast Modeling and Genetic Optimization Files

This submission includes example files associated with the Geothermal Operational Optimization using Machine Learning (GOOML) Big Kahuna fictional power plant, which uses synthetic data to model a fictional power plant. A forecast was produced using the GOOML data model framework and fictional input data, and a genetic optimization is included which determines optimal flash plant parameters. The inputs and outputs associated with the forecast and genetic optimization are included. The input and output files consist of data, configuration files, and plots. A link to the Physics-Guided Neural Networks (phygnn) GitHub repository is also included, which augments a traditional neural network loss function with a generic loss term that can be used to guide the neural network to learn physical or theoretical constraints. phygnn is used by the GOOML framework to help integrate its machine learning models into the relevant physics and engineering applications. Note that the data included in this submission are intended to provide a demonstration of GOOML's capabilities. Additional files that have not been released to the public are needed for users to run these models and reproduce these results. Units can be found in the readme data resource.

15 GEOTHERMAL ENERGY↗

Tsunami Early Warning From Global Navigation Satellite System Data Using Convolutional Neural Networks

Abstract We investigate the potential of using Global Navigation Satellite System (GNSS) observations to directly forecast full tsunami waveforms in real time. We train convolutional neural networks to use less than 9 min of GNSS data to forecast the full tsunami waveforms over 6 hr at select locations, and obtain accurate forecasts on a test data set. Our training and test data consists of synthetic earthquakes and associated GNSS data generated for the Cascadia Subduction Zone using the MudPy software, and corresponding tsunami waveforms in Puget Sound computed using GeoClaw. We use the same suite of synthetic earthquakes and waveforms as in earlier work where tsunami waveforms were used for forecasting, and provide a comparison. We also explore varying the number of GNSS stations, their locations, and their observation durations.

Rim, Donsub↗

Integrated Approach to Ancillary PV Component Reliability Assessment (Final Report)

In this project, we have established a nondestructive, generalized methodology that (1) fuses rich field data with advanced ML for proactive reliability forecasting, (2) dramatically reduces experimental iterations via synthetic dataset generation, and (3) achieves unprecedented regression precision in both anomaly detection and component-level degradation assessment—paving the way for truly predictive maintenance of grid-tied PV inverters under diverse outdoor conditions.

14 SOLAR ENERGY↗

Stochastic Price Generation for Evaluating Wholesale Electricity Market Bidding Strategies

This work presents a novel method for generating electricity price scenarios from statistical properties of past electricity prices using a hybrid statistical and reduced-form stochastic model. Previous work in applying stochastic differential equations (SDE) to model electricity prices has focused on daily average prices. To extend stochastic price generation methods to hourly or sub-hourly pricing, we address several weaknesses in the state-of-the-art: (1) we replace the mean-reversion component of the SDE with an ARIMA process that is better able to characterize the daily and weekly trends; (2) we extend the price-spike, or jump process to account for conditional probabilities of price spikes occurring in consecutive time steps by replacing the traditional Poisson process for modeling jumps with a generalized point process model inspired by brain neuron models; and (3) we replace the traditional method of estimating spike intensity with empirical variance with a Markov process based on observed price spike intensity transitions. The method is demonstrated with electricity prices from the US ERCOT market and a use-case example is provided for bidding an energy storage unit into the day-ahead and real-time energy markets of ERCOT using stochastic optimization methods. Results show that the the synthetic price model out performs a (naive) persistence forecast model by resulting in 24% to 47% more in profits over 168 simulated days.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

BuildingsBench: A Benchmark for Universal Building Load Forecasting [SWR-23-51]

The residential and commercial building stock in the United States is responsible for a significant percentage of energy consumption and greenhouse gas emissions. Electrification of end-uses, as well as decarbonizing the electrical grid through renewable energy sources such as solar and wind, constitutes the pathway to zero-emission buildings. Forecasting day-ahead building energy consumption is an integral part of this solution. Currently, specialized forecasting models are hand-made for each individual building, which is time-consuming, expensive, and leads to duplicated efforts. BuildingsBench is a Python software framework for training and comparing generalized machine learning models for universal building load forecasting. This challenge tasks a single foundational model to generalize its forecasts for a wide variety of buildings, across geographic regions, building types, weather patterns, and more. This software provide code for pre-training such models and subsequently evaluating their performance on a suite of hundreds of diverse real and synthetic buildings. BuildingsBench is a platform for: - Large-scale pretraining with the synthetic Buildings-900K dataset for short-term load forecasting (STLF). Buildings-900K is statistically representative of the entire U.S. building stock and is extracted from the NREL End-Use Load Profiles database. - Benchmarking on two tasks evaluating generalization: zero-shot STLF and transfer learning for STLF. We provide an index-based PyTorch Dataset for large-scale pretraining, easy data loading for multiple real building energy consumption datasets as PyTorch Tensors or Pandas DataFrames, simple (persistence) to advanced (transformer) baselines, metrics management, and more.

Emami, Patrick↗

State-of-the-Art Techniques for Large-Scale Stochastic Unit Commitment

Recent advances in deterministic unit commitment, both formulaic and algorithmic, along with modern algorithmic approaches for stochastic programming, have enabled the solution of stochastic unit commitment problems with hundreds of scenarios on large-scale transmission networks. In this presentation, we will give an overview of these methods, including lazy transmission constraint generation, lower-bounding techniques, and heuristics, all of which can be executed in concert with customized decomposition approaches for optimization under uncertainty. We demonstrate the effectiveness of these techniques on the TAMU Texas7K synthetic transmission network, leveraging realistic high-resolution forecasts based on NREL renewable resource availability data. The software leveraged for these demonstrations is available via the open-source software packages EGRET (for electrical grid optimization) and mpi-sppy (for optimization under uncertainty).

27 ARPA - Advanced Research Projects Agency-Energy↗