Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bad data detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

WISP: Watching grid Infrastructure Stealthily through Proxies (Final Technical Report)

The complex interdependencies of cyber systems (sensors and communications), physical grids and associated electricity market operations make protecting electric power grids a significant challenge. The energy sector is constantly under new, targeted, advanced and dangerous cyber-attacks that have the potential to result in the loss of human life. These threats are further exacerbated by our need to modernize the grid. One focus of cyber security research in smart grids is the securing of the SCADA system through advanced intrusion detection systems (IDS) and bad data detection algorithms in state estimation. These methods either require full knowledge of the system topology and parameters or fail to understand the physical behaviors under attack. WISP (Watching grid Infrastructure Stealthily through Proxies) is designed to provide additional protection to the power grid using only publicly available data. In particular, WISP exploits the spatio-temporal nature of the real time locational marginal prices (LMPs), in conjunction with other information such as bids, weather, outages and load data to analyze anomalous power pricing behaviors and then correlate those observations to localize regions of interest and identify potential cyber events. WISP is non-intrusive as the tool is deployed as a service in the Cloud or on premise and provides reliable information to system operators for enhanced situational awareness, without impeding energy delivery functions. The WISP technology comprises three modules: the data-driven anomaly detection core, the vulnerability and risk analysis and the root cause analysis. The data-driven anomaly detection core performs the tasks of feature selection, anomaly detection and attack region localization. The vulnerability and risk analysis module provides system level information of the vulnerable variables and times, assisting the operators in selecting monitoring and protection nodes. The root cause analysis module takes the detection results and identifies potential operational conditions that contribute to the detected anomalies. In Phase I, we have demonstrated the feasibility and effectiveness of WISP. We developed a realistic electricity market simulator capable of generating normal and attack market data under various operational conditions. We developed a series of cyber-attack detection and analysis algorithms and evaluated them under multiple data sources. Finally, we integrated all modules into an end-to-end software, providing functions for data management, data analytics and visualization. Specifically, we have achieved: (i) real-time data acceptance from external utility interfaces with >99% acceptance rate; (ii) high performance anomaly detection algorithms with >98% detection accuracy and <0.1% false alarm rate; and (iii) ultra-low computing delay <50 milliseconds. Additionally, our team developed algorithms to identify the vulnerable variables in electricity market operations and root cause analysis functions to identify major contributors to the price spikes. These ancillary modules are necessary when deploying WISP in real world industry environment. In Phase II, we have demonstrated the effectiveness of WISP software on realistic largescale power systems. We performed red team testing for the Phase I WISP software and identified software vulnerabilities and implemented corresponding mitigation solutions. We adapted the electricity market simulator for the Texas synthetic 2000-bus system and generated datasets for the false data injection attacks. We created database and visualization interfaces for the Texas system and the ISO New England system. We performed software optimization in terms of operation efficiency, computing speed and detection accuracy. Finally, we tested the software on the Texas system and the ISO New England system and evaluated the detection performance. Overall, we achieved above 89% detection rate, below 3% false alarm rate and below 37 seconds of end-to-end detection delay.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Moving-horizon false data injection attack design against cyber–physical systems

Systematic attack design is essential to understanding the vulnerabilities of cyber–physical systems (CPSs), to better design for resiliency. In particular, false data injection attacks (FDIAs) are well-known and have been shown to be capable of bypassing bad data detection (BDD) while causing targeted biases in resulting state estimates. However, their effectiveness against moving horizon estimators (MHE) is not well understood. In fact, this paper shows that conventional FDIAs are generally ineffective against MHE. One of the main reasons is that the moving window renders the static FDIA recursively infeasible. Here, this paper proposes a new attack methodology, moving-horizon FDIA (MH-FDIA), by considering both the performance of historical attacks and the current system’s status. Theoretical guarantees for successful attack generation and recursive feasibility are given. Numerical simulations on the IEEE-14 bus system further validate the theoretical claims and show that the proposed MH-FDIA outperforms state-of-the-art counterparts in both stealthiness and effectiveness. In addition, an experiment on a path-tracking control system of an autonomous vehicle shows the feasibility of the MH-FDIA in real-world nonlinear systems.

42 ENGINEERING↗

Anomaly Detection in Power System State Estimation: Review and New Directions

Foundational and state-of-the-art anomaly-detection methods through power system state estimation are reviewed. Traditional components for bad data detection, such as chi-square testing, residual-based methods, and hypothesis testing, are discussed to explain the motivations for recent anomaly-detection methods given the increasing complexity of power grids, energy management systems, and cyber-threats. In particular, state estimation anomaly detection based on data-driven quickest-change detection and artificial intelligence are discussed, and directions for research are suggested with particular emphasis on considerations of the future smart grid.

42 ENGINEERING↗

Detection of Stealthy False Data Injection Attacks in Unobservable Distribution Networks

In this paper, a composite scheme is proposed for detecting stealthy data manipulation attacks on distribution system which is unobservable with standard least squares based state estimators. This technique has three stages where the process of data imputation, voltage phasor estimation and the bad data detection are carried out in a systematic manner. The proposed approach is then integrated with moving target defense strategies which perturbs the network parameters to reveal stealthy false data injection attacks. The proposed approach is tested is validated on a three-phase, unbalanced 37-node distribution system and its results are presented. It is shown that the proposed approach has the ability to accurately detect the presence of FDI attacks using limited measurements (i.e., the test system is unobservable).

Rajasekaran, James K.↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

Distribution System State Estimation Using a Multiple Iteration Extended Kalman Filter Approach

To support the operation of modern distribution systems, operators require real-time visibility into system states. Due to a lack of measurements and unbalanced operation, the state estimation in distribution systems is challenging as compared to transmission systems. This paper proposes the utilization of a Multiple Iteration - Extended Kalman Filter based approach for the distribution system state estimation. This modified version of the baseline extended Kalman filter iterates over the update step multiple times thereby reducing the estimation error. The proposed algorithm along with the auxiliary algorithms such as bad data detection is integrated into a co-simulation environment. Case studies show that the proposed state estimation method can result in a lesser estimation error as compared to the baseline approach.

Bhatti, Bilal Ahmad↗

Robust Distribution State Estimation for Reliable Locational Marginal Pricing under Cyber-Attacks

Here this paper examines the impact of false data injection (FDI) cyber-attacks on distribution system state estimation (DSSE) and the resulting distribution locational marginal price (DLMP) in power markets. Two robust high-breakdown regression estimators, namely S- and MM- estimators, are implemented to provide resistance against FDI attacks targeting measurements and grid topology, creating leverage points. The introduced estimators are compared to the weighted least squares (WLS) with a bad data detection and rejection module (BDD) and the robust Huber M-estimator. The proposed estimators are shown to be effective and compare favorably to both existing Huber M- and the WLS with BDD in the presence of topology FDI attacks. Both the S- and MM-estimators provide good performance in the case of clean and corrupted measurements. Their performance is comparable in this case to the Huber M- and the WLS, followed by a BDD module. The simulation considered a modified distribution IEEE 13 and 34-bus systems where the impact of FDI attack scenarios is shown on the state and the DLMP pricing in the presence of distributed Generation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

AI & Physics-Based Bad Command/Data Detection in Large Power Electronics Systems: Multi-Port Autonomous Reconfigurable Solar Power Plant (MARS)

Detection of bad data from measurement sensors and bad commands from control centers need to be carried out to avoid instabilities within large power electronics systems. Towards the same, in this paper, model and data driven methods are proposed to identify anomalies in measured data and commands received by large power electronics systems. The large power electronics system considered in this paper is a multi-port autonomous reconfigurable solar power plant (MARS), which consists of photovoltaic (PV) and energy storage systems (ESSs) that connect to high-voltage direct current (HVdc) system and transmission ac power grid. The proposed algorithms in the MARS power plant to detect bad data from measurements and bad commands from control centers are evaluated in simulations and hardware-in-the-loop (HIL) tests. Furthermore, it has been observed that the proposed algorithms are able to detect bad measurements and commands in all the use cases evaluated.

14 SOLAR ENERGY↗

A Graph Convolutional Network for Active Distribution System Anomaly Detection Considering Measurement Spatial-Temporal Correlations

The accuracy of distribution system state estimation may be significantly impacted by the existence of bad measure-ments and unexpected topology errors. This paper proposes a data-driven Graph Convolutional Network (GCN) for anomaly detection, including bad measurements and topology change events. Compared to many existing machine learning approaches, the proposed approach embeds both spatial-temporal measure-ment correlations, which allows us to detect and distinguish different anomalies. Numerical results carried out on the IEEE 37-node system demonstrate that the proposed-based method can obtain high accuracy in detecting bad data and topology changes as compared to other approaches, even in the presence of high PV penetrations.

active distribution system↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Prediction of Power Measurements Using Adaptive Filters

With the advent of smart grid concept, Internet of Things (IoT) and the deployment of smart meters, the cyberattack threats on power networks have increased due to the use of communication systems that can be accessed by adversaries. Attackers will have the ability to manipulate the outcomes of smart meters which in turn influence the core application of Energy Management System 9EMS): State Estimation (SE). Bad data analytic tools may fail to detect some attacks into measurements. Meanwhile, Machine Learning (ML) solutions have been proposed for detecting False Data Injection (FDI) attacks. However, there is a lack of ML time-series solutions presented in the state-of-the-art that is yet to be complex. In signal processing, time-series solutions do not only consider the signal, but also the statistics of the signal over time. Therefore, in this paper, a machine learning for time-series solutions is presented as an application to model the measurements of the power grid that are used in SE. The presented model takes into account adaptive linear and non-linear filters: Finite Impulse Response (FIR), and Infinite Impulse Response (IIR). The presented models are implemented and performed on the IEEE-118 bus system. The results indicate the advantage of applying those filters over the state-of-the-art machine learning solutions.

Hamad, Khaled↗

Anomaly Detection and Mitigation in FACTS-based Wide-Area Voltage Control Systems using Machine Learning

With the increasing deployment of Flexible AC Transmission System (FACTS) devices in wide-area voltage control systems (WAVCS) for achieving improved voltage stability of bulk power systems, the possibility for cyber attacks on these systems is also increasing. Successful stealthy cyber attacks that are difficult to detect by traditional informational technology (IT)-based cybersecurity solutions or threshold-based bad data detectors can lead to a voltage collapse in power grid. This paper presents the testbed-based attacks implementation and real-time evaluation of machine learning (ML) algorithm for detecting and mitigating stealthy cyber attacks on FACTS-based WAVCS on a hardware-in-the-loop (HIL) testbed. Initially, we discuss the implementation of a fuzzy logic controller (FLC) that controls a Static VAR Compensator (SVC) device deployed in a two-area four-machine Kundur power system for improving transient voltage stability. Later, the ML-based Anomaly Detection and Mitigation (ADM) system is implemented on the cyber-physical HIL testbed to detect and mitigate various stealthy cyber attacks, which are injected in real-time over the wide-area network (WAN). The experimental results show accurate and effective performance of ADM system in detecting and mitigating anomalies while keeping the grid stable and within the system operating limits, as defined by the North America Electric Reliability Corporation (NERC).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Semi-Supervised Learning Method for the Identification of Bad Exposures in Large Imaging Surveys

As the data volume of astronomical imaging surveys rapidly increases, traditional methods for image anomaly detection, such as visual inspection by human experts, are becoming impractical. We introduce a machine-learning-based approach to detect poor-quality exposures in large imaging surveys, with a focus on the DECam Legacy Survey (DECaLS) in regions of low extinction (i.e., E ( B − V ) < 0.04 ). Our semi-supervised pipeline integrates a vision transformer (ViT), trained via self-supervised learning (SSL), with a k-Nearest Neighbor (kNN) classifier. We train and validate our pipeline using a small set of labeled exposures observed by surveys with the Dark Energy Camera (DECam). A clustering-space analysis of where our pipeline places images labeled in good and bad categories suggests that our approach can efficiently and accurately determine the quality of exposures. Applied to new imaging being reduced for DECaLS Data Release 11, our pipeline identifies 780 problematic exposures, which we subsequently verify through visual inspection. Being highly efficient and adaptable, our method offers a scalable solution for quality control in other large imaging surveys.

Luo, Yufeng (ORCID:0000000246230683)↗

High-Fidelity Dataset Generation for Sensor Anomalies in Power Grids using Hardware-in-the-Loop Testbed

Sensor anomalies in power grids can have significant impacts on the operation of the grid due to the increased reliance of the grid operation on data-driven applications. However, there is a lack of datasets that accurately capture these anomalies as many of the anomalies go undetected using the current bad data detectors. High-fidelity labeled datasets are essential for developing robust applications that can detect and mitigate the impacts of anomalies. In this paper, we propose a hardware-in-the-loop testbed model that can emulate the grid behavior with high-fidelity. This testbed is used to inject anomalies at various levels in the grid architecture and generate labeled datasets. These high-fidelity datasets can be used for development and validation of data-driven applications for detection and mitigation of anomalies in grids and other cyber-physical systems.

Hyder, Burhan↗

Multi-area parameter error identification for large power systems

Power grid model parameters may contain errors due to various reasons. Detecting and correcting parameter errors typically requires significant computational effort due to the size and complexity of the parameter database. While the normalized Lagrange multiplier (NLM) method can effectively detect, identify and correct parameter errors, its computational burden could rapidly grow with increasing system size. This paper addresses this issue by proposing a multi-area parameter error identification method. Each area has its own outlier detection tool for detecting the incorrect parameters and measurements within the area. On the other hand, due to the reduced redundancy at area boundaries, parameter errors on branches incident to boundary buses may not be detected. Such errors are subsequently detected by a coordination level estimator completing the system-wide parameter detection procedure. In conclusion, performance of the developed method is demonstrated using the IEEE 118-bus and 2000-bus Texas synthetic systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Robust Event Classification Using Imperfect Real-world PMU Data

Here, this paper studies robust event classification using imperfect real-world phasor measurement unit (PMU) data. By analyzing the real-world PMU data, we find it is challenging to directly use this dataset for event classifiers due to the low data quality observed in PMU measurements and event logs. To address these challenges, we develop a novel machine learning framework for training robust event classifiers, which consists of three main steps: data preprocessing, fine-grained event data extraction, and feature engineering. Specifically, the data preprocessing step addresses the data quality issues of PMU measurements (e.g., bad data and missing data); in the fine-grained event data extraction step, a model-free event detection method is developed to accurately localize the events from the inaccurate event timestamps in the event logs; and the feature engineering step constructs the event features based on the patterns of different event types, in order to improve the performance and the interpretability of the event classifiers. Based on the proposed framework, we develop a workflow for event classification using the real-world PMU data streaming into the system in real time. Using the proposed framework, robust event classifiers can be efficiently trained based on many off-the-shelf lightweight machine learning models. Numerical experiments using the real-world dataset from the Western Interconnection of the U.S power transmission grid show that the event classifiers trained under the proposed framework can achieve high classification accuracy while being robust against low-quality data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗