Search NASA⌕ Search

SEARCH · Search NASA

Results for “physics of faulting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Emulation and detection of physical faults and cyber-attacks on building energy systems through real-time hardware-in-the-loop experiments

The increasing use of remote or mobile access, integrated wearable technologies, data exchange, and cloud-based data analytics in modern smart buildings is steering the building industry towards open communication technologies. The increased connectivity and accessibility could lead to more cyber-attacks in smart buildings. On the other hand, physical faults (e.g., HVAC -heating, ventilation, and air-conditioning faults) may have similar adverse impacts as those from the cyber-attacks on building energy systems, such as occupant discomfort, energy wastage, and equipment downtime. However, current physical behavior-based anomaly detection methods fail to differentiate between cyber-attacks and physical faults in building energy systems. Moreover, the challenge in collecting real-world threat data with ground truth has led researchers to rely on numerical models with user-defined assumptions, which may not accurately reflect real-world conditions due to the lack of in-situ experimental datasets. To address these challenges and gaps, this paper presents a flexible hardware-in-the-loop (HIL) testbed for generating cyber-attack and physical fault datasets and demonstrating threat detection algorithms in a real building automation system (BAS) environment. This testbed combines hardware (i.e., real BAS with local HVAC controllers and a physical network) with software (i.e., high-fidelity models to represent behaviors of building envelope and HVAC energy systems), enabling emulations of realistic threats. Five HIL experiments, including one baseline without any threats, two with physical faults, and two with cyber-attacks, were conducted to generate datasets containing detailed network traffic and system states. A joint classification framework, incorporating a network analyzer and a physical HVAC fault detector, was proposed to automatically detect cyber-physical abnormalities on BAS at both the network and the physical HVAC levels. The network analyzer comprises a conditional random fields (CRF) based command validator and a statistics-based detection strategy. The fault detector employs a weather and schedule-based pattern matching and feature-based principal component analysis (WPM-FPCA) method. Evaluation of the classification using four metrics from the multi-class confusion matrix revealed an average accuracy of 90.2%, recall of 89.7%, precision of 88.5% and F1-score of 89.2%. Finally, these results demonstrate that the proposed joint classification framework can effectively differentiate between specific types of cyber-attacks (e.g., device reinitialization attack, network Denial-of-Service attack) and physical faults (e.g., air handling unit operational fault, cooling coil valve stuck) in real time for improved building energy management.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Experimental analysis of computer system dependability

This paper reviews an area which has evolved over the past 15 years: experimental analysis of computer system dependability. Methodologies and advances are discussed for three basic approaches used in the area: simulated fault injection, physical fault injection, and measurement-based analysis. The three approaches are suited, respectively, to dependability evaluation in the three phases of a system's life: design phase, prototype phase, and operational phase. Before the discussion of these phases, several statistical techniques used in the area are introduced. For each phase, a classification of research methods or study topics is outlined, followed by discussion of these methods or topics as well as representative studies. The statistical techniques introduced include the estimation of parameters and confidence intervals, probability distribution characterization, and several multivariate analysis methods. Importance sampling, a statistical technique used to accelerate Monte Carlo simulation, is also introduced. The discussion of simulated fault injection covers electrical-level, logic-level, and function-level fault injection methods as well as representative simulation environments such as FOCUS and DEPEND. The discussion of physical fault injection covers hardware, software, and radiation fault injection methods as well as several software and hybrid tools including FIAT, FERARI, HYBRID, and FINE. The discussion of measurement-based analysis covers measurement and data processing techniques, basic error characterization, dependency analysis, Markov reward modeling, software-dependability, and fault diagnosis. The discussion involves several important issues studies in the area, including fault models, fast simulation techniques, workload/failure dependency, correlated failures, and software fault tolerance.

Iyer, Ravishankar, K.↗

Using Markov Models of Fault Growth Physics and Environmental Stresses to Optimize Control Actions

A generalized Markov chain representation of fault dynamics is presented for the case that available modeling of fault growth physics and future environmental stresses can be represented by two independent stochastic process models. A contrived but representatively challenging example will be presented and analyzed, in which uncertainty in the modeling of fault growth physics is represented by a uniformly distributed dice throwing process, and a discrete random walk is used to represent uncertain modeling of future exogenous loading demands to be placed on the system. A finite horizon dynamic programming algorithm is used to solve for an optimal control policy over a finite time window for the case that stochastic models representing physics of failure and future environmental stresses are known, and the states of both stochastic processes are observable by implemented control routines. The fundamental limitations of optimization performed in the presence of uncertain modeling information are examined by comparing the outcomes obtained from simulations of an optimizing control policy with the outcomes that would be achievable if all modeling uncertainties were removed from the system.

Bole, Brian↗

A hardware-in-the-loop (HIL) testbed for cyber-physical energy systems in smart commercial buildings

In recent years, there has been a growing trend toward the development of smart buildings that rely on cyber-physical systems (CPS) to optimize occupant comfort, safety, and energy efficiency. To ensure the reliable and efficient operation of CPS with designed control strategies, it is important to evaluate their performance under various scenarios before deploying them in the real world. This is where a Hardware-in-the-loop (HIL) testbed designed for studying sensor and control-related studies in smart buildings can be highly valuable. With the growing threat of cyber-attacks and physical faults targeting smart buildings, it is essential to ensure the security of building operations. A HIL testbed can emulate cyber-attack and physical fault scenarios, allowing researchers to develop and test threat detection and mitigation algorithms. This enables researchers to identify potential issues and optimize the algorithms in a safe and controlled environment before they are deployed in real-world settings, reducing the risk of failures that can negatively impact occupant comfort, safety, and energy efficiency. Therefore, this paper developed a HIL testbed designed for cyber-physical energy systems (e.g. buildings automation system (BAS)) in smart commercial buildings. The HIL testbed is comprised of a real-time building and Heating, Ventilation, and Air-Conditioning (HVAC) emulator using Modelica-based dynamic models, a set of BAS controllers, and a BAS computer server. The data generation capability of the HIL testbed is demonstrated by tracking normal and faulty operating data in the BAS, as well as monitoring detailed network traffic in the local BAS network. Here, this study further demonstrates the HIL testbed’s capability by conducting case studies on real-time physical fault and cyber-attack experiments using a Department of Energy (DOE) prototype commercial building. It is anticipated that the fully functional HIL testbed will be utilized for a variety of sensor and control-related studies, including but not limited to testing, developing, validating of different HVAC control strategies, fault detection & diagnosis, energy monitoring and analysis, cyber security study, etc.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

DRS: Derivational Reasoning System

The high reliability requirements for airborne systems requires fault-tolerant architectures to address failures in the presence of physical faults, and the elimination of design flaws during the specification and validation phase of the design cycle. Although much progress has been made in developing methods to address physical faults, design flaws remain a serious problem. Formal methods provides a mathematical basis for removing design flaws from digital systems. DRS (Derivational Reasoning System) is a formal design tool based on advanced research in mathematical modeling and formal synthesis. The system implements a basic design algebra for synthesizing digital circuit descriptions from high level functional specifications. DRS incorporates an executable specification language, a set of correctness preserving transformations, verification interface, and a logic synthesis interface, making it a powerful tool for realizing hardware from abstract specifications. DRS integrates recent advances in transformational reasoning, automated theorem proving and high-level CAD synthesis systems in order to provide enhanced reliability in designs with reduced time and cost.

Bose, Bhaskar↗

CYDRES: CYber Defense and REsilient System for securing grid-interactive efficient buildings

Smart buildings, especially Grid-interactive Efficient Buildings (GEBs), suffer from cyber-attacks and physical faults due to the integration of a large number of sensors and controls, connected devices, and associated communication networks. This study demonstrated a real-time advanced building resilient platform, called CYber Defense and REsilient System (CYDRES), which is deployable for existing and emerging Building Automation Systems (BASs). CYDRES aims to empower GEBs with cyber-attack-immune capabilities through multi-layer prevention and adaptation mechanisms to monitor, detect, and respond to cyber-attacks and physical operational faults. CYDRES is demonstrated through real-time experiments in a Hardware-in-the-Loop (HIL) testbed.

Building automation system, Cyber-attacks, Physica↗

Physics-constrained fault diagnosis framework for monitoring a multi-component thermal hydraulic system

A method for diagnosing faults includes receiving a system description of a thermal hydraulic system, the system description indicating a plurality of components and sensors. The method also includes constructing, based on physical conservations laws and using the system description, a plurality of physics-based models for the plurality of components, each of the plurality of physics-based models including unknown parameters. The method further includes receiving historical measurements and calibrating the physics-based models by calculating the unknown parameters of each of the physics-based models using the historical measurements to produce calibrated models. The method also includes receiving sensor measurements of the sensors, and calculating residuals corresponding to differences between measurements predicted by the plurality of calibrated models and the sensor measurements. The method also includes determining, based on the calculated residuals, a fault of a component or a sensor, and generating an alert indicating the fault.

Nguyen, Tat Nghia↗

Physics-constrained fault diagnosis framework for monitoring a standalone component of a thermal hydraulic system

A method for diagnosing faults includes receiving a description of a component of a thermal hydraulic system, where the description also indicates one or more sensors of the component. The method also includes constructing, based on a physical conservation law and using the description, a physics-based model describing operation of the component, the physics-based model including one or more unknown parameters. The method further includes calibrating the physics-based model by calculating the one or more unknown parameters using historical measurements to produce a calibrated model. Further, the method includes receiving sensor measurements captured by the one or more sensors, and calculating residuals corresponding to differences between measurements predicted by the calibrated model and the sensor measurements. The method also includes determining, based on the calculated residuals, a fault of the component or of a sensor of the one or more sensors, and generating an alert indicating the fault.

Nguyen, Tat Nghia↗

Modeling and Evaluation of Cyber-Attacks on Grid-Interactive Efficient Buildings

Grid-interactive efficient buildings (GEBs) are not only exposed to passive threats (e.g., physical faults) but also active threats such as cyber-attacks launched on the network-based control systems. The impact of cyber-attacks on GEB operation are not yet fully understood, especially as regards the performance of grid services. To quantify the consequences of cyber-attacks on GEBs, this paper proposes a modeling and simulation framework that includes different cyber-attack models and key performance indexes to quantify the performance of GEB operation under cyber-attacks. The framework is numerically demonstrated to model and evaluate cyber-attacks such as data intrusion attacks and Denial-of-Service attacks on a typical medium-sized office building that uses the BACnet/IP protocol for communication networks. Simulation results show that, while different types of attacks could compromise the building systems to different extents, attacks via the remote control of a chiller yield the most significant consequences on a building system’s operation, including both the building service and the grid service. It is also noted that a cyber-attack impacts the building systems during the attack period as well as the post-attack period, which suggests that both periods should be considered to fully evaluate the consequences of a cyber-attack.

Fu, Yanyang↗

Simulation-based reasoning about the physical propagation of fault effects

The research described deals with the effects of faults on complex physical systems, with particular emphasis on aircraft and spacecraft systems. Given that a malfunction has occurred and been diagnosed, the goal is to determine how that fault will propagate to other subsystems, and what the effects will be on vehicle functionality. In particular, the use of qualitative spatial simulation to determine the physical propagation of fault effects in 3-D space is described.

Feyock, Stefan↗

Estimating the distribution of fault latency in a digital processor

Presented is a statistical approach to measuring fault latency in a digital processor. The method relies on the use of physical fault injection where the duration of the fault injection can be controlled. Although a specific fault's latency period is never directly measured, the method indirectly determines the distribution of fault latency.

Ellis, Erik L.↗

A New On-Line Diagnosis Protocol for the SPIDER Family of Byzantine Fault Tolerant Architectures

This paper presents the formal verification of a new protocol for online distributed diagnosis for the SPIDER family of architectures. An instance of the Scalable Processor-Independent Design for Electromagnetic Resilience (SPIDER) architecture consists of a collection of processing elements communicating over a Reliable Optical Bus (ROBUS). The ROBUS is a specialized fault-tolerant device that guarantees Interactive Consistency, Distributed Diagnosis (Group Membership), and Synchronization in the presence of a bounded number of physical faults. Formal verification of the original SPIDER diagnosis protocol provided a detailed understanding that led to the discovery of a significantly more efficient protocol. The original protocol was adapted from the formally verified protocol used in the MAFT architecture. It required O(N) message exchanges per defendant to correctly diagnose failures in a system with N nodes. The new protocol achieves the same diagnostic fidelity, but only requires O(1) exchanges per defendant. This paper presents this new diagnosis protocol and a formal proof of its correctness using PVS.

Geser, Alfons↗

AI for Earthquake Physics

The core LANL program sponsored by Office of Science, Basic Energy Science, Chemical Sciences, Geosciences, and Biosciences (DOE-BES-CSGB) and led by PI Johnson aims to research earthquake faults to advance fault physics and earthquake hazards. All work completed is required to be made publicly available through publications and open-source codes supporting the published results. All routines are/will-be written in open source python and applied to publicly available data sets. These routines will format data from input into models, develop and test modeling frameworks for the problems addressed, and produce figures applicable to peer-reviewed manuscripts. All work is reviewed for Los Alamos Unlimited Release before submitting to a journal. This summary encompasses recently completed work and work to be complete for the duration of the program.

Johnson, Christopher↗

Implementation of an experimental fault-tolerant memory system

The experimental fault-tolerant memory system described in this paper has been designed to enable the modular addition of spares, to validate the theoretical fault-secure and self-testing properties of the translator/corrector, to provide a basis for experiments using the new testing and correction processes for recovery, and to determine the practicality of such systems. The hardware design and implementation are described, together with methods of fault insertion. The hardware/software interface, including a restricted single error correction/double error detection (SEC/DED) code, is specified. Procedures are carefully described which, (1) test for specified physical faults, (2) ensure that single error corrections are not miscorrections due to triple faults, and (3) enable recovery from double errors.

Carter, W. C.↗

Development of a Unified Taxonomy for HVAC System Faults

Detecting and diagnosing HVAC faults is critical for maintaining building operation performance, reducing energy waste, and ensuring indoor comfort. An increasing deployment of commercial fault detection and diagnostics (FDD) software tools in commercial buildings in the past decade has significantly increased buildings’ operational reliability and reduced energy consumption. A massive amount of data has been generated by the FDD software tools. However, efficiently utilizing FDD data for ‘big data’ analytics, algorithm improvement, and other data-driven applications is challenging because the format and naming conventions of those data are very customized, unstructured, and hard to interpret. This paper presents the development of a unified taxonomy for HVAC faults. A taxonomy is an orderly classification of HVAC faults according to their characteristics and causal relations. The taxonomy includes fault categorization, physical hierarchy, fault library, relation model, and naming/tagging scheme. The taxonomy employs both a physical hierarchy of HVAC equipment and a cause-effect relationship model to reveal the root causes of faults in HVAC systems. A structured and standardized vocabulary library is developed to increase data representability and interpretability. The developed fault taxonomy can be used for HVAC system ‘big data’ analytics such as HVAC system fault prevalence analysis or the development of an HVAC FDD software standard. A common type of HVAC equipment-packaged rooftop unit (RTU) is used as an example to demonstrate the application of the developed fault taxonomy. Two RTU FDD software tools are used to show that after mapping FDD data according to the taxonomy, the meta-analysis of the multiple FDD reports is possible and efficient.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

General linear codes for fault-tolerant matrix operations on processor arrays

Various checksum codes have been suggested for fault-tolerant matrix computations on processor arrays. Use of these codes is limited due to potential roundoff and overflow errors. Numerical errors may also be misconstrued as errors due to physical faults in the system. In this a set of linear codes is identified which can be used for fault-tolerant matrix operations such as matrix addition, multiplication, transposition, and LU-decomposition, with minimum numerical error. Encoding schemes are given for some of the example codes which fall under the general set of codes. With the help of experiments, a rule of thumb for the selection of a particular code for a given application is derived.

Nair, V. S. S.↗

Hardware-in-the-Loop Testbed for Cyber-Physical Security of Photovoltaic Farms

In the last decades, modem grids with distributed energy resources, such as photovoltaic (PV) farms, are increasingly vulnerable to cyber-attacks that seriously affect the stability and performance of the power system. While cyber-physical security of smart grids is extensively studied, most of the existing work focuses on the grid level and neglects the modeling and features of device-level power electronics converters (PECs). Furthermore, establishing a high-fidelity simulation testbed that can simulate harmonic frequencies of the PV farm is in urgent need. In this paper, a high-fidelity and real-time hardware-in- the-loop testbed is built to simulate the harmonics of power electronics converters for cyber-physical security of PEC-enabled PV farms. Based on this testbed, the impact of typical cyber-attacks and physical faults on the PV converter can be analyzed, thus providing a foundation for cyber-attack detection, root cause diagnosis, and resilient control to mitigate the adverse effects of cyber-attacks.

14 SOLAR ENERGY↗