Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Cybersecurity Center for Offshore Wind Energy (Final Project Report)

This project establishes a Cybersecurity Center for Offshore Wind Energy with the objective of designing and operating a cyber-physical testbed for wind energy farms (WEFs) that enables comprehensive cybersecurity research. The testbed incorporates a Supervisory Control and Data Acquisition (SCADA) system connected to turbine models via industrial-grade programmable logic controllers (PLCs) and remote terminal units (RTUs). It supports side-channel data acquisition, implementation and analysis of various cyberattack scenarios, and development of attack detection, mitigation, and best-practice guidance tailored to wind energy systems. During the project, the team expanded the number and fidelity of mathematical turbine models (MTMs), integrated these models with SCADA infrastructure, and deployed a scaled physical turbine and associated sensors. High-resolution operational and side-channel data streams were collected and used to refine machine-learning (ML)-based attack detection systems and to extend the WindCRAFT framework to multi-turbine threat scenarios. The project demonstrated a realistic, scalable environment for evaluating cyber threats, validated attack detection approaches using enriched datasets, and identified new multi-turbine and inter-turbine communication attack vectors. The resulting testbed, models, and security mechanisms provide a foundation for ongoing R&D and deployment of cyber-resilient offshore wind energy systems.

17 WIND ENERGY↗

Probing the matter-dominated expansion with multi-redshift Lyman-$α$ BAO from DESI DR2

We present a multi-redshift Baryon Acoustic Oscillations (BAO) analysis of the DESI Data Release 2 (DR2) Lyman-$α$ (Ly$α$) forest, splitting the forest auto-correlation and its cross-correlation with quasars into three redshift bins. We obtain BAO measurements at effective redshifts $z_{\rm eff} = 2.13$, $2.40$, and $2.81$ with $\sim2.0$--$2.5\%$ precision per bin in the radial and transverse directions, corresponding to $\sim1.1$--$1.2\%$ precision for the isotropic BAO measurement. Using the same data products and modeling framework as the DESI DR2 Ly$α$ BAO analysis, we validate the pipeline on $400$ synthetic datasets and find unbiased BAO recovery with well-calibrated uncertainties. The measurements show an increase in the isotropic dilation parameter $D_V/r_d$ from $30.26\pm0.39$ to $32.22\pm0.47$ and in the Alcock-Paczyński parameter $D_M/D_H$ from $3.96\pm0.15$ to $5.63^{+0.22}_{-0.24}$. The Hubble distance $D_H/r_d$ decreases from $9.40\pm0.20$ to $7.22\pm0.17$, providing a direct measurement of the expansion history consistent with $Λ$CDM and the expected matter-dominated scaling, with $H(z)\propto(1+z)^n$ giving $n=1.34\pm0.16$. The redshift split also provides a self-consistent measurement of clustering evolution: the Ly$α$ forest bias evolves as $(1+z)^γ$ with $γ_α=3.05\pm0.16$, the RSD parameter has a redshift evolution described by $γ_β=-0.97\pm0.26$, and the quasar bias evolves with $γ_Q=1.56\pm0.23$, consistent with independent quasar clustering measurements. Combining these three-bin BAO measurements with DESI DR2 galaxy and quasar BAO measurements yields cosmological constraints consistent with the single-bin Ly$α$ BAO analysis in flat $Λ$CDM and $w_0w_a$CDM and improves curvature constraints by $\sim12\%$ in $Λ$CDM$+Ω_\mathrm{K}$.

Herrera-Alcantar, Hiram K. [Paris, Inst. Astrophys↗

Manufacturing Facility Inventory National Dataset (M-FIND)

This asset provides a high-fidelity, validated inventory of manufacturing facilities across the United States, filling a critical gap in publicly available industrial data. By integrating and cross-referencing thirteen distinct data sources, this dataset moves beyond the limitations of single-source registries to provide a harmonized list that includes precise geographic coordinates, industrial subsector designations, and—crucially—parcel-level spatial boundaries.

Billings, Blake [ORNL] (ORCID:0000000186021600)↗

Diesel Fuel Consumption in Prominent U.S. Open-Pit Mines: Site-Level Estimates

This report presents a comprehensive framework for estimating diesel fuel consumption and prices at open-pit mines in the United States. The framework includes transparent methods for calculating site-level diesel energy use when direct reporting is unavailable, and a structured confidence evaluation for each method. The framework is demonstrated to estimate current diesel consumption at 21 open-pit mines in the United States. Initial findings support ongoing efforts to strengthen the competitiveness and security of the U.S. industrial base by supporting data-driven supply chain analysis and decision-making, improved transparency in mining sector energy use, and targeted deployment of energy innovation and cost-reduction strategies. Future updates to the dataset—coupled with expanded data transparency and method validation—will help ensure that the findings remain relevant as the sector continues to evolve.

02 PETROLEUM↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Designing resilient IoT and Edge Computing with federated tinyML

The rapid growth of the Internet of Things (IoT) and Edge Computing (EC) has brought significant conveniences to modern society but has also greatly expanded the cyber attack surfaces, particularly as these technologies are being increasingly integrated into critical systems such as power grids, healthcare, and smart homes. Here, to improve IoT/EC’s cybersecurity posture, we leveraged Artificial Intelligence (AI) and Machine Learning (ML) by employing tinyML to monitor voluminous IoT data for cyber threats while addressing devices’ resource constraints, and utilizing Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. Building on our three-layer architecture combining tinyML and FL to enhance autonomous cyber attack detection, this paper demonstrated that the architecture improves detection accuracy, reduces resource consumption, and enables lightweight, secure IoT device monitoring. These results were validated using the public N-BaIoT dataset as well as real IoT network traffic data collected under multiple attack scenarios from our testbeds. Additionally, we introduced an enhanced FL methodology with a novel preprocessing stage, including federated feature selection and global preprocessor construction, to address IoT/EC data heterogeneity. We developed a physical IoT testbed for attack simulations and data collection, implemented a tinyML-powered detector for realistic model validation, and also built a virtual testbed for scalable evaluations of FL models across diverse network environments.

Cognitive cyber↗

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

Validation of the DESI DR2 Ly$α$ forest full-shape analysis

We present the validation of the Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) Lyman-$α$ (Ly$α$) forest full-shape analysis. This analysis combines three-dimensional Ly$α$ forest auto-correlations and cross-correlations with quasars to extract information from both the baryon acoustic oscillation (BAO) feature and the broadband clustering signal, with primary emphasis on the Alcock-Paczynski (AP) measurement. Compared to the DESI DR1 analysis, the DR2 validation uses substantially larger and more realistic mock datasets, including CoLoRe 2LPT and AbacusSummit Ly$α$ forest simulations. The modeling framework is also improved through analytic marginalization over small scales ($<10$$h^{-1}$Mpc) and the impact of ultraviolet background fluctuations. The validation program was completed prior to unblinding and defines quantitative requirements for the cosmological parameters of interest, which are evaluated using hundreds of mock realizations. We further test the analysis through independent fits to the auto- and cross-correlations, multiple catalog splits, and a broad suite of analysis and modeling variations applied to both mocks and blinded observational data. We find that the BAO and AP parameters satisfy all validation requirements and remain stable across all tests. In contrast, mock studies reveal a significant bias in the inferred growth-rate parameter $fσ_8$, leading us to exclude this measurement from the final analysis. The consistency across mocks, data splits, and robustness tests demonstrates that the DR2 Ly$α$ full-shape analysis provides a reliable and substantially improved broadband AP measurement over previous Ly$α$ forest studies.

Herbold, M. [Chicago U., KICP; Ohio State U.] (ORC↗

Demonstration of gold nanorod systems for enhanced total efficiency: Experimental and numerical analysis

This study investigates the photothermal performance of gold nanorods engineered to exhibit longitudinal plasmon resonances at 695 nm, 780 nm, and 970 nm. The work combines synthesis, structural characterization, extinction measurements, numerical modeling, and controlled temperature experiments to quantify how nanorod geometry, resonance tuning, concentration, and chamber shape jointly influence heat generation. Transmission electron microscopy confirms that increasing nanorod aspect ratio systematically shifts the longitudinal plasmon peak toward the near-infrared region. Extinction measurements show strong agreement with theoretical predictions, validating the numerical model across two independent datasets. Three chamber geometries were tested under laser excitation at 640 nm, 808 nm, and 980 nm: an ascending stepped base, a flat base, and a descending stepped base. Without nanorods, the ascending geometry produced the highest efficiency due to enhanced natural convection. After introducing gold nanorods, all geometries exhibited substantial thermal enhancement, with total efficiencies exceeding 20%. The strongest improvement was obtained for nanorods resonant at 780 nm with a mass concentration of 4.6 mg/mL implemented on the descending stepped-base geometry. This performance resulted from the combined effect of spectral overlapping with the 808 nm laser, the highest nanorod concentration, and localized heat accumulation that intensified buoyancy-driven flow. The findings demonstrate that total efficiency is governed by a synergistic interplay between optical resonance, nanoparticle concentration, and macroscopic chamber design, revealing the system-level coupling between nanoscale plasmonic absorption and macroscale heat-transfer phenomena. The results provide a validated framework for tuning nanoscale plasmonic absorbers and optimizing thermal systems for applications requiring efficient light-to-heat conversion.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

The Zooplankton International Geospatial (ZIG) dataset: A global repository of spatiotemporal freshwater zooplankton community composition data to support ecological research

Zooplankton play critical roles in aquatic ecosystem function and food webs. Nevertheless, global syntheses of their abundance and community dynamics are challenging due to methodological differences across monitoring programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we assembled, curated, validated, and harmonized the Zooplankton International Geospatial (ZIG) dataset, which includes co-located and contemporaneous zooplankton, water chemistry, and limnological data from 307 lakes and reservoirs. ZIG includes waterbodies from each major lake thermal region and range in size from 0.8-2,805,8600 hectares. Temporal coverage for individual waterbodies ranges between 1-60 years of data (median = 4 years) with sampling from once annually to weekly. ZIG is publicly available and can be used to understand freshwater biodiversity change and its drivers at unprecedented scales, and we consider it to be a cornerstone for future investigations of freshwater biology, chemistry, and ecology.

Figary, Stephanie [Cornell University, Ithaca, NY]↗

Datasets and U-Net Model for "A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma"

This dataset has results and the model associated with the publication Ciulla et al., (2024). It contains a U-Net semantic segmentation model (unet_model.h5) and associated code implemented in tensorflow 2.0 for the model training and identification of oil and gas well symbols in USGS historical topographic maps (HTMC). Given a quadrangle map (7.5 minutes), downloadable at this url: https://ngmdb.usgs.gov/topoview/, and a list of coordinates of the documented wells present in the area, the model returns the coordinates of oil and gas symbols in the HTMC maps. For reproducibility of our workflow, we provide a sample map in California and the documented well locations for the entire State of California (CalGEM_AllWells_20231128.csv) downloaded from https://www.conservation.ca.gov/calgem/maps/Pages/GISMapping2.aspx. Additionally, the locations of 1,301 potential undocumented orphaned wells identified using our deep learning framework or the counties of Los Angeles and Kern in California, and Osage and Oklahoma in Oklahoma are provided in the file found_potential_UOWs.zip. The results of the visual inspection of satellite imagery in Osage County is in the file visible_potential_UOWs.zip. The dataset also includes a custom tool to validate the detected symbols in the HTMC maps (vetting_tool.py). More details about the methodology can be found in the associated paper: Ciulla, F., Santos, A., Jordan, P., Kneafsey, T., Biraud, S.C., and Varadharajan, C. (2024) A Deep Learning Based Framework to Identify Undocumented Orphaned Oil and Gas Wells from Historical Maps: a Case Study for California and Oklahoma. Accepted for publication in Environmental Science and Technology. The geographical coordinates provided correspond to the locations of potential undocumented orphaned oil and gas wells (UOWs) extracted from historical maps. The actual presence of wells need to be confirmed with on-the-ground investigations. For your safety, do not attempt to visit or investigate these sites without appropriate safety training, proper equipment, and authorization from local authorities. Approaching these well sites without proper personal protective equipment (PPE) may pose significant health and safety risks. Oil and gas wells can emit hazardous gasses including methane, which is flammable, odorless and colorless, as well as hydrogen sulfide, which can be fatal even at low concentrations. Additionally, there may be unstable ground near the wellhead that may collapse around the wellbore. This dataset was prepared as an account of work sponsored by the United States Government. While this document is believed to contain correct information, neither the United States Government nor any agency thereof, nor the Regents of the University of California, nor any of their employees, makes any warranty, express or implied, or assumes any legal responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by its trade name, trademark, manufacturer, or otherwise, does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or the Regents of the University of California. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof or the Regents of the University of California.

Artificial Intelligence↗

Burst pressure models and validations for thick-walled pipelines containing corrosion defects

Corrosion is one major threat to pipeline integrity. Over the past decades, many corrosion models have been developed for determining the remaining strength of corroded pipelines, including ASME B31.G, Modified B31.G, LPC, PCORRC and their modified models. All these corrosion models are applicable only to large diameter, thin-walled pipelines with a diameter to wall thickness ratio D/t ≥ 20. In practice, many pipelines have a small diameter and thick wall with a D/t ratio < 20, and thus an adequate corrosion model is needed for assessing remaining strength for corroded thick-walled pipelines. This paper briefly reviews the theoretical burst pressure models for defect-free thin and thick-walled pipelines and four representative corrosion assessment models for thin-walled corroded pipelines. On this basis, two modified corrosion models are proposed to thick-walled pipelines in terms of the average shear stress yield theory. To verify the proposed corrosion models, comprehensive validations are performed. Numerical validations include the elastic-plastic finite element analysis to determine burst pressure for pipelines without and with corrosion defects and the model evaluation using a large dataset of available FEA results of burst pressure for machined defects. Experimental validations include a set of burst pressure tests for defect-free thick-walled pipes with different thicknesses and the model evaluation using one large burst dataset for machined defects with flat bottoms and another large dataset for real corrosion defects with curved river bottom profiles. Both numerical and experimental validations show that the proposed corrosion models can more accurately predict the remaining strength for corroded thin and thick-walled pipelines.

Pipeline↗

WTK-LED: The WIND Toolkit Long-Term Ensemble Dataset

To satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have, we aimed at developing an updated version of the meteorological WIND Toolkit, named WIND Toolkit Long-term Ensemble Dataset (WTK-LED), which is a meteorological dataset providing time series every 5 min and 2 km, including model uncertainty of wind speed at every modeling grid point so that users are provided with a range of possible wind speeds every 2 km. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: 1) Numerical simulations covering the continental United States, Alaska, and Hawaii, with high-resolution data being available for 3 years (2018-2020). 2) Climate simulations from Argonne National Laboratories covering the North American continent, including Alaska, Canada, and most of Mexico and the Caribbean Islands. These simulations complement the new WTK-LED to offer a 4-km dataset covering 20 years, from 2001-2020. 3) Specific long-term,high-resolution offshore simulations have been conducted separately for the US coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind data set. This report focuses on a description of the land-based WTK-LED for CONUS, Hawaii, and Alaska, for the 3-year 2-km/5-min dataset and the 20-year 4-km/hourly dataset, as well as the uncertainty quantification method. We also provide limited validation results. Based on our results to date, we suggest use cases and applications for each dataset of the WTK-LED.

17 WIND ENERGY↗

Application of Machine Learning and Data Augmentation Algorithms in the Discovery of Metal Hydrides for Hydrogen Storage

The development of efficient and sustainable hydrogen storage materials is a key challenge for realizing hydrogen as a clean and flexible energy carrier. Among various options, metal hydrides offer high volumetric storage density and operational safety, yet their application is limited by thermodynamic, kinetic, and compositional constraints. In this work, we investigate the potential of machine learning (ML) to predict key thermodynamic properties—equilibrium plateau pressure, enthalpy, and entropy of hydride formation—based solely on alloy composition using Magpie-generated descriptors. We significantly expand an existing experimental dataset from ~400 to 806 entries and assess the impact of dataset size and data augmentation, using the PADRE algorithm, on model performance. Models including Support Vector Machines and Gradient Boosted Random Forests were trained and optimized via grid search and cross-validation. Results show a marked improvement in predictive accuracy with increased dataset size, while data augmentation benefits are limited to smaller datasets and do not improve accuracy in underrepresented pressure regimes. Furthermore, clustering and cross-validation analyses highlight the limited generalizability of models across different material classes, though high accuracy is achieved when training and testing within a single hydride family (e.g., AB2). The study demonstrates the viability and limitations of ML for accelerating hydride discovery, emphasizing the importance of dataset diversity and representation for robust property prediction.

augmentation↗

Characterizing in-stream turbulent flow for tidal energy converter siting in Cook Inlet, Alaska

Cook Inlet in Alaska is the most promising location for tidal energy development in the U.S. due to its significant tidal range of approximately 10 meters and high volume flux. The inlet's unique geometry and flow characteristics make it the most energetic tidal stream in the nation, with GW-scale potential energy capacity. With the growing interest in tidal energy converter (TEC) deployment in this area, we implemented a regional-scale, 3D hydrodynamic modeling framework to predict tidal current and turbulence characteristics that can assist TEC designers and project managers. We validated the model results extensively using various datasets collected with bottom-mounted acoustic Doppler current profilers and velocimeters. The comparison between the model outputs and observational data highlighted the effectiveness of the 3D FVCOM model and the Mellor-Yamada Level 2.5 Turbulence Model in accurately assessing macro-scale kinetic energy, turbulence intensity, and the production and dissipation rates at a prospective TEC site. Using two months of model simulation data, we examined the channel cross-section for TEC deployment, focusing on undisturbed power density and macro-scale turbulent properties. Further, our findings indicate that understanding the turbulence characteristics and flow properties can enhance Stage I/II resource characterization by identifying optimal locations for TECs and their layouts within the channel. Furthermore, we demonstrated that TEC designers can utilize macro-scale turbulence data from 3D coastal models as boundary conditions for other turbulence models, allowing for a more detailed resolution of the turbulence structure at TEC siting locations. Ultimately, this work emphasizes the importance of estimating flow and turbulence conditions in energetic systems to understand turbulent sites better and improve resource characterization.

16 TIDAL AND WAVE POWER↗

DECADE+ DES Y3 Weak Lensing Mass Map: A 13,000 deg $\^{} 2$ View of Cosmic Structure from 270 Million Galaxies

We present the largest galaxy weak lensing mass map of the late-time Universe, reconstructed from 270 million galaxies in the DECADE and DES Year 3 datasets, covering 13,000 square degrees. We validate the map through systematic tests against observational conditions (depth, seeing, etc.), finding the map is statistically consistent with no contamination. The large area covered by the mass map makes it a well-suited tool for cosmological analyses, cross-correlation studies and the identification of large-scale structure features. We demonstrate its potential by detecting cosmic filaments directly from the mass map for the first time and validating them through their association with galaxy clusters selected using the Sunyaev-Zeldovich effect from Planck and ACT DR6.

Gatti, M.↗

SO(3)-invariant PCA with application to molecular data

Principal component analysis (PCA) is a fundamental technique for dimensionality reduction and denoising; however, its application to three-dimensional data with arbitrary orientations -- common in structural biology -- presents significant challenges. A naive approach requires augmenting the dataset with many rotated copies of each sample, incurring prohibitive computational costs. In this paper, we extend PCA to 3D volumetric datasets with unknown orientations by developing an efficient and principled framework for SO(3)-invariant PCA that implicitly accounts for all rotations without explicit data augmentation. By exploiting underlying algebraic structure, we demonstrate that the computation involves only the square root of the total number of covariance entries, resulting in a substantial reduction in complexity. We validate the method on real-world molecular datasets, demonstrating its effectiveness and opening up new possibilities for large-scale, high-dimensional reconstruction problems.

Fraiman, Michael [Tel Aviv Univ., Tel Aviv (Israel↗

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan↗