Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Functional Data Analysis in Wearable Body Sensor Networks

Improving response time of indirect room-size calorimeters is still an outstanding problem in metabolic research. Accurate estimates of instantaneous rates of gaseous exchange require numerical differentiation of measured gaseousgas concentrations. We propose a new method to estimate the instantaneous gaseousgas exchange rates in indirect calorimetry. In contrast to the previously developed techniques, the method addresses the problem of differentiation of gaseous concentrations as an ill-posed problem. By applying the method of regularization, the problem of differentiation is converted into a well-posed problem resulting in smooth and consistent gaseous exchange rates. The validity of the method is tested on a large dataset of calorimeter experiments which included 313 human experiments along with 231 alcohol combustion experiments. It is demonstrated that the method is able to reliably differentiate between the “unphysiological” process of alcohol combustion and physiological variations produced by human metabolism. The method also allowed unraveling the previously unreported relative kinetics of O2 consumption and Respiratory Quotient (RQ) in humans. It was found that the kinetics of oxidative fuel selection lags behind the energy expenditure in humans exhibiting some sort of oxidative inertia. The time lag varies from 2-3 min up to 30 min, depending on particular individual. No such lag was found in alcohol combustion experiments. In addition to the relative kinetics of substrate oxidation, two statistical indexes reflecting variability of minute-by-minute RQ were estimated. The indexes were the RQ’s standard deviation and RQ’s first-order derivative. Both indexes showed statistically significant difference between human experiments and alcohol combustion experiments. We conclude that the proposed method can consistently extract physiologically-relevant information from noisy calorimetry data and the aforesaid information can provide additional insights into the mechanism of metabolic fuel selection in humans.

54 ENVIRONMENTAL SCIENCES↗

Detecting Neutrons in MicroBooNE

A significant challenge in measurements of neutrino oscillations is reconstructing the incoming neutrino energies. While modern fully-active tracking calorimeters such as liquid argon time projection chambers in principle allow the measurement of all final state particles above some detection threshold, undetected neutrons remain a considerable source of missing energy with little to no data constraining their production rates and kinematics. We present the first demonstration of tagging neutrino-induced neutrons in liquid argon time projection chambers using secondary protons emitted from neutron-argon interactions in the MicroBooNE detector. We describe the method developed to identify neutrino-induced neutrons and demonstrate its performance using neutrons produced in muon-neutrino charged current interactions. The method is validated using a small subset of MicroBooNE's total dataset. The selection yields a sample with $60\%$ of selected tracks corresponding to neutron-induced secondary protons.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Functional Data Analysis in Wearable Body Sensor Networks

Improving response time of indirect room-size calorimeters is still an outstanding problem in metabolic research. Accurate estimates of instantaneous rates of gaseous exchange require numerical differentiation of measured gaseousgas concentrations. We propose a new method to estimate the instantaneous gaseousgas exchange rates in indirect calorimetry. In contrast to the previously developed techniques, the method addresses the problem of differentiation of gaseous concentrations as an ill-posed problem. By applying the method of regularization, the problem of differentiation is converted into a well-posed problem resulting in smooth and consistent gaseous exchange rates. The validity of the method is tested on a large dataset of calorimeter experiments which included 313 human experiments along with 231 alcohol combustion experiments. It is demonstrated that the method is able to reliably differentiate between the “unphysiological” process of alcohol combustion and physiological variations produced by human metabolism. The method also allowed unraveling the previously unreported relative kinetics of O2 consumption and Respiratory Quotient (RQ) in humans. It was found that the kinetics of oxidative fuel selection lags behind the energy expenditure in humans exhibiting some sort of oxidative inertia. The time lag varies from 2-3 min up to 30 min, depending on particular individual. No such lag was found in alcohol combustion experiments. In addition to the relative kinetics of substrate oxidation, two statistical indexes reflecting variability of minute-by-minute RQ were estimated. The indexes were the RQ’s standard deviation and RQ’s first-order derivative. Both indexes showed statistically significant difference between human experiments and alcohol combustion experiments. We conclude that the proposed method can consistently extract physiologically-relevant information from noisy calorimetry data and the aforesaid information can provide additional insights into the mechanism of metabolic fuel selection in humans.

60 - APPLIED LIFE SCIENCES↗

The DECADE cosmic shear project III: validation of analysis pipeline using spatially inhomogeneous data

We present the pipeline for the cosmic shear analysis of the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog consisting of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. The catalog derives from a large number of disparate observing programs and is therefore more inhomogeneous across the sky compared to existing lensing surveys. First, we use simulated data-vectors to show the sensitivity of our constraints to different analysis choices in our inference pipeline, including sensitivity to residual systematics. Next we use simulations to validate our covariance modeling for inhomogeneous datasets. Finally, we show that our choices in the end-to-end cosmic shear pipeline are robust against inhomogeneities in the survey, by extracting relative shifts in the cosmology constraints across different subsets of the footprint/catalog and showing they are all consistent within 1σ to 2σ. This is done for forty-six subsets of the data and is carried out in a fully consistent manner: for each subset of the data, we re-derive the photometric redshift estimates, shear calibrations, survey transfer functions, the data vector, measurement covariance, and finally, the cosmological constraints. Our results show that existing analysis methods for weak lensing cosmology can be fairly resilient towards inhomogeneous datasets. This also motivates exploring a wider range of image data for pursuing such cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗

IM3 Projected U.S. Western Interconnection Grid Stress Dataset

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from GO WEST and TEP) for Integrated Multisector, Multiscale Modeling (IM3) Phase 2 simulations across eight different scenarios for the U.S. Western Interconnection through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States from a set of Thermodynamic Global Warming (TGW) simulations. These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 GO WEST is an open-source power grid modeling framework for the U.S. Western Interconnection, which allows users to tailor the model depending on their research study and science questions. It covers 28 balancing authorities (BAs) and 12 states in U.S. Western Interconnection. GO WEST allows users to select different number of nodes and come up with a simplified network by utilizing 10,000 nodal topology of the U.S. Western Interconnection (ACTIVSg10k). Users can select different number of nodes, mathematical formulations (linear programming vs. mixed-integer linear programming), transmission line limit scaling factors, and hurdle rate scaling factors. GO WEST offers a unit commitment and economic dispatch (UC/ED) module to simulate grid operations on an hourly scale. In this sense, users can calibrate and validate their model versions by comparing model outputs to historical datasets. TEP is an open-source transmission capacity expansion model, built on the GO WEST framework. It utilizes linear programming to optimize transmission capacity addition investment on existing lines within the GO WEST framework. The TEP model only increases the thermal capacity of existing transmission lines and does not add new lines to the system, which leaves the topology preserved. In order to use TEP model, users need to create scenarios with the GO WEST framework. Please refer to README file for a detailed description of the dataset including individual files and references.

Capacity Expansion Model↗

Fidelity-preserving enhancement of ptychography with foundational text-to-image models

Ptychographic phase retrieval enables high-resolution imaging of complex samples but often suffers from artifacts such as grid pathology and multislice crosstalk, which degrade reconstructed images. We propose a plug-and-play (PnP) framework that integrates physics model-based phase retrieval with text-guided image editing using foundational diffusion models. By employing the alternating direction method of multipliers, our approach ensures consensus between data fidelity and artifact removal subproblems, maintaining physical consistency while enhancing image quality. Artifact removal is achieved using a text-guided diffusion image editing method (LEDITS++) with a pre-trained foundational diffusion model, allowing users to specify artifacts for removal in natural language. Demonstrations on simulated and experimental datasets show significant improvements in artifact suppression and structural fidelity, validated by metrics such as peak signal-to-noise ratio and diffraction pattern consistency. This work highlights the combination of text-guided generative models and model-based phase retrieval algorithms as a transferable and fidelity-preserving method for high-quality diffraction imaging.

image editing↗

Autonomie Simulation Datasets in Support of U.S. DOT-NHTSA Advanced Vehicle Technology Research

Understanding how new vehicle technologies affect fuel economy and energy use is critical to the regulatory work performed by the U.S. Department of Transportation’s National Highway Traffic Safety Administration (NHTSA), which sets Corporate Average Fuel Economy (CAFE) standards under the Energy Policy and Conservation Act of 1975. In order to support this work, Argonne National Laboratory uses Autonomie, a full-vehicle simulation tool, to evaluate advanced powertrain architectures and their effects on vehicle energy consumption and performance. A wide range of vehicle classes has been assessed (i.e., internal combustion engine vehicles, hybrid electric vehicles, plug-in hybrid electric vehicles, battery-electric vehicles, and fuel cell electric vehicles), as well as the effects of various technology improvements such as lightweighting, aerodynamic refinements, and low-rolling-resistance tires. Simulations have been run across multiple drive cycles to capture fuel and electricity use under realistic operating conditions. The resulting datasets include detailed vehicle-level results, model assumptions, and validation reports, all of which have been made publicly available through NHTSA in support of the 2023 notice of proposed rulemaking covering light-duty vehicles for model years 2027 to 2035. These data are critical to stakeholders working in fuel economy regulation, vehicle technology assessment, and energy policy analysis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Cooperative Transmission Expansion Planning Experiment Data and Results

GO WEST is an open-source power grid modeling framework for U.S. Western Interconnection, which allows users to tailor the model depending on their research study and science questions. It covers 28 balancing authorities (BA) and 12 states in U.S. Western Interconnection. GO WEST allows users to select different number of nodes and come up with a simplified network by utilizing 10,000 nodal topology of U.S. Western Interconnection created by Texas A&M University. Users can try and select different number of nodes, mathematical formulations (linear programming vs. mixed-integer linear programming), transmission line limit scaling factors, and hurdle rate scaling factors. GO WEST offers a unit commitment and economic dispatch (UC/ED) module to simulate grid operations on an hourly scale. In this sense, users can calibrate and validate their model versions by comparing model outputs to historical datasets. TEP is an open-source transmission capacity expansion model, built on GO WEST framework. It utilizes linear programming to optimize transmission capacity addition investment on existing lines within GO WEST framework. In this sense, TEP model only increases the thermal capacity of existing transmission lines and does not add new lines to the system, which leaves the topology preserved. TEP minimizes the total cost of the system which comprises the operational cost of satisfying electricity demand (i.e., generation cost), cost of loss of load (i.e., unserved energy), cost of power flow, and cost of new transmission capacity additions (i.e., investment cost). In order to use TEP model, users need to create scenarios with GO WEST framework. In this analysis, outputs from several models are used to create future inputs to GO WEST and TEP models, including GCAM-USA, TELL, CERF and reV. This dataset includes experiment inputs and outputs from three different transmission expansion scenarios (cooperative, intermediate, and individual) for 2019 and 2059. For 2019, a base scenario to illustrate the default (i.e., historical) power grid operations is also included. This study utilizes rcp45hotter_ssp3 scenario from a previous version of GCAM-USA simulations. Sources of the shapefiles in supplementary data are HIFLD Open and U.S. Energy Atlas. Please see the README file for a detailed description of the main and supplementary data.

Capacity Expansion Model↗

Measuring the neutrino-oxygen neutral current quasielastic cross section using the accelerator neutrino neutron interaction experiment

The Accelerator Neutrino Neutron Interaction Experiment (ANNIE) is a 26-ton gadolinium-doped water Cherenkov detector located on-axis to Fermilab’s Booster Neutrino Beam (BNB). ANNIE is uniquely positioned to perform high-statistics measurements of neutrino-nucleus interactions in water, benefiting from a large neutrino flux due to a short (100-meter) baseline. A central focus of ANNIE’s physics program is the measurement of both charged current (CC) and neutral current (NC) cross sections on water, including neutral current quasielastic (NCQE) and CC-inclusive channels. The NCQE measurement is particularly critical for constraining uncertainties in rare-event searches such as the Diffuse Supernova Neutrino Background (DSNB), where atmospheric $\nu$NCQE interactions constitute a significant and poorly constrained background. This dissertation presents a measurement of the flux-averaged neutrino-oxygen neutral current quasielastic ($\nu$NCQE) cross section using $2.573 \times 10^{20}$~POT of BNB exposure from the 2022 and 2023 beam years. The $\nu$NCQE interaction is identified through the primary $\gamma$-rays produced by nuclear de-excitation of the residual $^{15}$N$^*$ or $^{15}$O$^*$ nucleus following nucleon knockout from $^{16}$O. A dedicated Monte Carlo (MC) re-tuning campaign was conducted using an americium-beryllium (AmBe) calibration source, Michel electrons from stopped muons, and throughgoing dirt muons originating upstream of the detector. This multi-sample approach provided a wide-ranging $\mathcal{O}(\text{MeV})$--$\mathcal{O}(\text{GeV})$ dataset for tuning the simulated detector response, which was subsequently validated against AmBe neutron and Michel electron data for use in the $\nu$NCQE analysis. A dedicated laser calibration campaign was carried out to reduce timing uncertainties across the PMT system, enabling reconstruction of the BNB bunch substructure with sufficient resolution to serve as a background rejection tool. By selecting events in-time with individual neutrino bunches, beam-correlated $\nu$NCQE events are separated from diffuse and accelerator-induced backgrounds, notably skyshine neutrons and externally-originating events, that would otherwise dominate traditional charge-based selections within a small-scale, surface-level, short-baseline detector. A data-driven estimation of the skyshine neutron and external background rates was performed and incorporated into the systematic uncertainty budget. The flux-averaged $\nu$NCQE cross section on oxygen is measured to be $1.57 \pm 0.06\,(\text{stat.})$ $^{+0.91}_{-0.67}\,(\text{syst.})$ $\times 10^{-38}\ \text{cm}^{2}$. A full systematic budget is constructed by propagating uncertainties in the secondary hadronic interaction modeling, background cross section normalizations, detector response, neutrino flux, and the primary $\gamma$-ray emission probabilities from oxygen nuclear de-excitation. An idealized de-excitation model, constructed from existing measurements in the literature is developed to benchmark the predictions of the \textsc{GENIE} event generator. A comparison reveals that \textsc{GENIE} systematically overpredicts the primary $\gamma$-ray emission probability from oxygen de-excitation by a factor of $1.49\times$ for $E_\gamma > 6$~MeV and $3.07\times$ in the $3$--$6$~MeV band. This comparison motivates the dominant systematic uncertainty in this analysis, where a conservative uncertainty of $^{+39.9\%}_{-0\%}$ on the primary $\gamma$-ray signal prediction is assigned. The ANNIE result is consistent with and complementary to existing flux-averaged $\nu$NCQE cross section measurements from T2K and Super-Kamiokande, providing an independent measurement with a different detector, neutrino beam, and analysis methodology. Looking ahead, an upgrade to the ANNIE DAQ infrastructure enabling continuous extended readout will allow a complementary $\nu$NCQE neutron multiplicity measurement, directly relevant to constraining the NCQE background in DSNB searches, competitive with the recent T2K measurement at SK-Gd. The planned Super-SANDI upgrade, deploying a large Water-based Liquid Scintillator (WbLS) volume, will further extend ANNIE's reach to hadronic final states and exclusive NC channels, and enable joint measurements with liquid argon detectors sharing the BNB beamline ahead of DUNE and Hyper-Kamiokande.

Doran, Steven [Iowa State U.]↗

Constraining the Stellar-to-Halo Mass Relation with Galaxy Clustering and Weak Lensing from DES Year 3 Data

We develop a framework to study the relation between the stellar mass of a galaxy and the total mass of its host dark matter halo using galaxy clustering and galaxy-galaxy lensing measurements. We model a wide range of scales, roughly from $\sim 100 \; {\rm kpc}$ to $\sim 100 \; {\rm Mpc}$, using a theoretical framework based on the Halo Occupation Distribution and data from Year 3 of the Dark Energy Survey (DES) dataset. The new advances of this work include: 1) the generation and validation of a new stellar mass-selected galaxy sample in the range of $\log M_\star/M_\odot \sim 9.6$ to $\sim 11.5$; 2) the joint-modeling framework of galaxy clustering and galaxy-galaxy lensing that is able to describe our stellar mass-selected sample deep into the 1-halo regime; and 3) stellar-to-halo mass relation (SHMR) constraints from this dataset. In general, our SHMR constraints agree well with existing literature with various weak lensing measurements. We constrain the free parameters in the SHMR functional form $\log M_\star (M_h) = \log(εM_1) + f\left[ \log\left( M_h / M_1 \right) \right] - f(0)$, with $f(x) \equiv -\log(10^{αx}+1) + δ[\log(1+\exp(x))]^γ/ [1+\exp(10^{-x})]$, to be $\log M_1 = 11.506^{+0.325}_{-0.404}$, $\log ε= -1.632^{+0.306}_{-0.181}$, $α= -1.638^{+0.108}_{-0.099}$, $γ= 0.596^{+0.251}_{-0.210}$ and $δ= 3.810^{+2.045}_{-1.811}$. The inferred average satellite fraction is within $\sim 5-35\%$ for our fiducial results and we do not see any clear trends with redshift or stellar mass. Furthermore, we find that the inferred average galaxy bias values follow the generally expected trends with stellar mass and redshift. Our study is the first SHMR in DES in this mass range, and we expect the stellar mass sample to be of general interest for other science cases.

Zacharegkas, G. [Argonne; Chicago U., KICP] (ORCID↗

Dataset of Generative AI Workload Power Profiles

This dataset provides a collection of high-resolution (5/10 Hz or every 0.2/0.1 seconds) power consumption profiles for generative artificial intelligence (GenAI) workloads executed on NLR's High Performance Computing (HPC) platform Kestrel. The dataset also includes examples of representative whole-facility power profiles generated using a bottom-up, event-driven, data center energy model . This dataset is designed to support research in energy modeling, infrastructure planning, energy system integration, and sustainability analysis for AI-driven computing systems. The dataset captures time-resolved electrical power measurements across a diverse set of configurations, including variations in job type (inference vs. training), workload (LLM vs. image generation), datasets, and number of compute nodes. Power traces are provided in a standardized format and include both raw/instantaneous and aggregated files. Each profile is accompanied by metadata describing workload parameters, enabling reproducibility and cross-study comparison. The dataset is intended for use in applications such as data center infrastructure planning, energy modeling, demand response and grid impact studies, and development and validation of system-level simulation tools. By making these workload-specific power profiles publicly available, this dataset aims to address the current lack of open, empirical energy data for generative AI systems and to facilitate transparent, reproducible research on the energy and environmental impacts of large-scale AI deployment. If you use this dataset, please cite the associated publication: Vercellino et al., “Measurement of Generative AI Workload Power Profiles for Whole-Facility Data Center Infrastructure Planning,” arXiv:2604.07345 (2026).

97 MATHEMATICS AND COMPUTING↗

Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets

This study addresses the challenge of statistically extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings. We investigate encoder-decoder-based generative models for nonlinear dimensionality reduction, focusing on disentangling low-dimensional latent variables corresponding to independent physical factors. Introducing Aux-VAE, a novel architecture within the classical Variational Autoencoder framework, we achieve disentanglement with minimal modifications to the standard VAE loss function by leveraging prior statistical knowledge through auxiliary variables. These variables guide the shaping of the latent space by aligning latent factors with learned auxiliary variables. We validate the efficacy of Aux-VAE through comparative assessments on multiple datasets, including astronomical simulations.

97 MATHEMATICS AND COMPUTING↗

Using multiple high-resolution datasets to benchmark the energy exascale earth system model (E3SM) for renewable resource assessment

The United States is accelerating its shift toward a renewable energy system. However, renewable resources, which harness energy from the Earth system, are susceptible to both present-day climate variability and future climate change. For example, variations in regional climate can alter renewable energy production patterns and site viability. The use of high-resolution climate model projections can therefore facilitate and may be critical to long-term planning of renewable energy investments. However, climate models must first be validated for renewable resource assessment. This research employs multiple high-spatiotemporal-resolution datasets to assess the capability of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 North American Regionally Refined Model (E3SMv2-NARRM) for predicting multi-year climatological values of solar and wind energy capacity factors in the continental U.S., with a focus on regional and seasonal variability. Present-day E3SMv2-NARRM simulations are compared with reported utility-scale production data obtained from the Energy Information Administration (EIA). In addition, E3SMv2-NARRM data are evaluated against non-climate benchmark models from the National Renewable Energy Laboratory, including the Wind Integration National Dataset Toolkit and the National Solar Radiation Database (NSRDB), as well as three wind energy datasets from PLUSWIND. Our analysis indicates that solar capacity factors from E3SM closely match those from the NSRDB dataset. However, both datasets tend to overestimate values by 10% in comparison to EIA data. Furthermore, biases in wind capacity factors within E3SM are notably pronounced in the West Coast regions, where the seasonal cycle diverges from EIA data.

Energy forecasting, Capacity factor, Renewable ene↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

High-Resolution South American Wind Resource Data Downscaled with Generative Machine Learning Conditioned on Near-Surface Observations

High-resolution historical wind data was developed for the entirety of South America using the innovative Super-Resolution for Renewable Resource Data (sup3r) machine learning framework. The publicly available Sup3rWind South America dataset represents a significant advancement in wind resource data generation, leveraging generative machine learning conditioned on near-surface observations from the Meteorological Assimilation Data Ingest System (MADIS) to efficiently and accurately downscale coarse reanalysis data from the European Centre for Medium-Range Weather Forecasts (ERA5). This approach produces fine-scale, spatially and temporally coherent wind and meteorological fields hundreds of times more computationally efficient than traditional numerical weather modeling methods, enabling access to high-fidelity wind information across both continental and offshore regions. Sup3rWind South America builds on the earlier Sup3rWind Ukraine dataset through improvements in model architecture and outputs conditioned on near-surface observation inputs. As with the Ukraine data release, this dataset includes wind speed, wind direction, temperature, relative humidity, and pressure at a horizontal resolution of ~2 km, representing a 15x spatial enhancement relative to the 31 km ERA5 grid. Wind speed and direction are provided at 5-minute resolution, a 12x temporal refinement compared to the hourly ERA5 data, while temperature, relative humidity, and pressure remain at hourly resolution. The data covers all years from 2005 to 2024. Before downscaling, ERA5 inputs were bias-corrected using long-term monthly means and a limited number of quality-controlled observations to align large-scale statistics with regional conditions. The resulting dataset is the first publicly available high-resolution timeseries wind record that provides full spatial coverage of South America. Model validation demonstrates strong agreement with observations across several statistical metrics, consistent with other state-of-the-art high-resolution wind resource datasets. The potential applications of Sup3rWind South America span renewable energy resource assessment, energy system modeling, and grid resilience analysis. The 20-year record and high spatial and temporal resolution support accurate estimation of long-term energy yield and the economic feasibility of potential wind development sites. Continuous coverage across both continental and offshore regions enables comprehensive site prospecting within exclusive economic zones. The 2 km, 5-minute resolution data provide the spatial and temporal variability required for power system simulation, operational planning, and regional risk assessments.

17 WIND ENERGY↗

A Neural Optimizer With Decision-Focused Learning for Optimal Energy Storage Operation

Here, this article introduces a neural optimizer-based framework for optimizing battery energy storage system (BESS) control for grid services, including demand charge and energy cost reduction. By leveraging decision-focused learning (DFL), the proposed framework ensures seamless integration and adaptation, significantly enhancing control performance. A patch time-series transformer is employed for peak load forecasting, incorporating aleatoric uncertainty quantification to account for forecasting uncertainties within the decision-making process. The framework utilizes a solver-in-the-loop approach to generate optimal BESS actions, which are then used to train the neural optimizer-based agent. By co-optimizing both BESS operational modes and output power within the NN, the system achieves improved performance and robustness. After initial training, the forecasting and control models are jointly fine-tuned to account for forecasting errors, further improving decision precision and efficiency through DFL. Case studies are performed to validate the performance of the framework using multiple real-world datasets, demonstrating superior performance in monthly peak load forecasting compared to state-of-the-art models. In addition, the results are compared against existing decision-making approaches. The results demonstrate a reduction in monthly peak forecasting error by approximately 15% across various performance measures and achieve an optimization gap for BESS operation that is about three times smaller compared to existing methods.

Kim, Hyeonjin [Pacific Northwest National Laborato↗

From Sim to Real: A Pipeline for Training and Deploying Traffic Smoothing Cruise Controllers

Designing and validating controllers for connected and automated vehicles to enhance traffic flow presents significant challenges, from the complexity of replicating real-world stop-and-go traffic dynamics in simulation, to the intricacies involved in transitioning from simulation to actual deployment. In this work, we present a full pipeline from data collection to controller deployment. Specifically, we collect 772 km of driving data from the I-24 in Tennessee, and use it to build a one-lane simulator, placing simulated vehicles behind real-world trajectories. Using policy-gradient methods with an asymmetric critic, we improve fuel efficiency by over 10% when simulating congested scenarios. Our comprehensive approach includes reinforcement learning for controller training, software verification, hardware validation and setup, and navigating various sim-to-real challenges. Furthermore, we analyze the controller's behavior and wave-smoothing properties, and deploy it on four Toyota Rav4’s in a real-world validation experiment on the I-24. Lastly, we release the driving dataset, the simulator and the trained controller, to enable future benchmarking and controller design.

42 ENGINEERING↗