Search NASA⌕ Search

SEARCH · Search NASA

Results for “unreliable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Robust Decentralized Learning Using ADMM With Unreliable Agents

Many signal processing and machine learning problems can be formulated as consensus optimization problems which can be solved efficiently via a cooperative multi-agent system. However, the agents in the system can be unreliable due to a variety of reasons: noise, faults and attacks. Providing erroneous updates leads the optimization process in a wrong direction, and degrades the performance of distributed machine learning algorithms. This paper considers the problem of decentralized learning using ADMM in the presence of unreliable agents. First, we rigorously analyze the effect of erroneous updates (in ADMM learning iterations) on the convergence behavior of the multi-agent system. We show that the algorithm linearly converges to a neighborhood of the optimal solution under certain conditions and characterize the neighborhood size analytically. Next, we provide guidelines for network design to achieve a faster convergence to the neighborhood. Here, we also provide conditions on the erroneous updates for exact convergence to the optimal solution. Finally, to mitigate the influence of unreliable agents, we propose ROAD , a robust variant of ADMM, and show its resilience to unreliable agents with an exact convergence to the optimum.

97 MATHEMATICS AND COMPUTING↗

Unreliability of two-band model analysis of magnetoresistivities in unveiling temperature-driven Lifshitz transition

Recently, anomalies in the temperature dependences of the carrier density and/or mobility derived from analysis of the magnetoresistivities using the conventional two-band model have been used to unveil intriguing temperature-induced Lifshitz transitions in various materials. For instance, two temperature-driven Lifshitz transitions were inferred to exist in the Dirac nodal-line semimetal ZrSiSe, based on two-band model analysis of the Hall magnetoconductivities where the second band exhibits a change in the carrier type from holes to electrons when the temperature decreases below T=106K and a dip is observed in the mobility vs temperature curve at T=80K. Here, in this study, we revisit the experiments and two-band model analysis on ZrSiSe. We show that the anomalies in the second band may be spurious because the first band dominates the Hall magnetoconductivities at T>80K, making the carrier type and mobility obtained for the second band from the two-band model analysis unreliable. That is, care must be taken in interpreting these anomalies as evidence for temperature-driven Lifshitz transitions. Our skepticism on the existence of such phase transitions in ZrSiSe is further supported by the validation of Kohler's rule for magnetoresistances for T≤180K. In this paper, we showcase potential issues in interpreting anomalies in the temperature dependence of the carrier density and mobility derived from the analysis of magnetoconductivities or magnetoresistivities using the conventional two-band model.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Semi-Markov Unreliability-Range Evaluator

Reconfigurable, fault-tolerant systems modeled. Semi-Markov unreliability-range evaluator (SURE) computer program is software tool for analysis of reliability of reconfigurable, fault-tolerant systems. Based on new method for computing death-state probabilities of semi-Markov model. Computes accurate upper and lower bounds on probability of failure of system. Written in PASCAL.

Butler, Ricky W.↗

The cost of unreliability in the avionics suite of a launch vehicle

The main objective of the Advanced Launch System (ALS) program is to realize a substantial reduction in recurring launch costs over present launch systems. A methodology is presented for assessing the impact of the reliability of the avionics suite on the recurring launch cost of the ALS. The methodology is illustrated by focusing on two architectures. The first is a distillation of a typical architecture for this type of application. The second is a simplified implementation of the Advanced Information Processing System (AIPS) technology developed at the Charles Stark Draper Laboratory, Inc. Both architectures utilize redundancy of hardware to tolerate faults. the key contributors to the cost of unreliability are identified and modeled utilizing Markov modeling techniques for each of the two architectures. The results are presented along with the more traditional costs for these avionics suites.

Rosch, Gene↗

Semi-Markov Unreliability Range Evaluator

Semi-Markov Unreliability Range Evaluator, SURE, computer program is software tool for analysis of reconfigurable, fault-tolerant systems. Traditional reliability analyses based on aggregates of fault-handling and fault-occurrence models. SURE provides efficient means for calculating accurate upper and lower bounds for probabilities of death states for large class of semi-Markov mathematical models, and not merely those reduced to critical-pair architectures.

Butler, Ricky W.↗

Algorithms for Multiple Fault Diagnosis With Unreliable Tests

In this paper, we consider the problem of constructing optimal and near-optimal multiple fault diagnosis (MFD) in bipartite systems with unreliable (imperfect) tests. It is known that exact computation of conditional probabilities for multiple fault diagnosis is NP-hard. The novel feature of our diagnostic algorithms is the use of Lagrangian relaxation and subgradient optimization methods to provide: (1) near optimal solutions for the MFD problem, and (2) upper bounds for an optimal branch-and-bound algorithm. The proposed method is illustrated using several examples. Computational results indicate that: (1) our algorithm has superior computational performance to the existing algorithms (approximately three orders of magnitude improvement), (2) the near optimal algorithm generates the most likely candidates with a very high accuracy, and (3) our algorithm can find the most likely candidates in systems with as many as 1000 faults.

Shakeri, Mojdeh↗

Preliminary Results for Using Uncertainty and Out-of-distribution Detection to Identify Unreliable Predictions.

As machine learning (ML) models are deployed into an ever-diversifying set of application spaces, ranging from self-driving cars to cybersecurity to climate modeling, the need to carefully evaluate model credibility becomes increasingly important. Uncertainty quantification (UQ) provides important information about the ability of a learned model to make sound predictions, often with respect to individual test cases. However, most UQ methods for ML are themselves data-driven and therefore susceptible to the same knowledge gaps as the models themselves. Specifically, UQ helps to identify points near decision boundaries where the models fit the data poorly, yet predictions can score as certain for points that are under-represented by the training data and thus out-of-distribution (OOD). One method for evaluating the quality of both ML models and their associated uncertainty estimates is out-of-distribution detection (OODD). We combine OODD with UQ to provide insights into the reliability of the individual predictions made by an ML model.

97 MATHEMATICS AND COMPUTING↗

The semi-Markov unreliability range evaluator program

The SURE program is a design/validation tool for ultrareliable computer system architectures. The system uses simple algebraic formulas to compute accurate upper and lower bounds for the death state probabilities of a large class of semi-Markov models. The mathematical formulas used in the program were derived from a mathematical theorem proven by Allan White under contract to NASA Langley Research Center. This mathematical theorem is discussed along with the user interface to the SURE program.

Butler, R. W.↗

Semi-Markov Unreliability Range Evaluator (SURE)

Analysis tool for reconfigurable, fault-tolerant systems, SURE provides efficient way to calculate accurate upper and lower bounds for death state probabilities for large class of semi-Markov models. Calculated bounds close enough for use in reliability studies of ultrareliable computer systems. Written in PASCAL for interactive execution and runs on DEC VAX computer under VMS.

Butler, R. W.↗

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence on a Time Limited Mission?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach applies to the case where the unit failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, the design then uses N redundant units, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is FN = F1N, and the needed N = LN(FN)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. The needed redundancy, N, to achieve the specified N unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - achieves the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some uncertainty, the confidence that the rate is not lower than the actual failure rate and the required reliability is not overestimated is about 50%. Adding more redundant units increases the confidence that the required reliability, FN, will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals have higher confidence in being achieved. Both the required reliability and confidence can be specified initially and the needed number of redundant units computed using the measured failure rate. The unit failure rate is determined by initial reliability growth testing to remove design errors and to better estimate the final constant failure rate. Reducing the failure rate and reducing its variance both reduce the number of redundant units needed for the required reliability and confidence. Since the total cost is the sum of the costs of the units and of the testing, there is an optimum test time that produces minimum cost.

Harry W. Jones↗

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach is developed for the case when the system failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, N redundant units can be used, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is Ffail = F1 N , and the needed redundancy N = LN(F)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. (When F1 = f * L << 1, F1 is the probability that a unit will fail during the mission. When F1 = f * L > 1, F1 is the expected number of failures during the mission.) The needed redundancy, N, to achieve the required N redundant unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - is equal to the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some probabilistic uncertainty, the actual failure rate will be randomly higher or lower. This means that the reliability of the N redundant systems will be overestimated about half the time. Adding more redundant units increases the confidence that the required reliability will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals will be achieved with higher confidence. Both the desired reliability and confidence can be specified as initial requirements and the needed number of redundant units estimated using the measured failure rate.

Redundancy↗

APIS: Honeybee Foraging Task Assignment for Use in Uncertain and Unreliable Environments

Multiagent Cyber-Physical-Human (CPH) systems in realistic environments operate under uncertain conditions. Communication among agents, aimed at reducing the uncertainty, is itself subject to uncertainty. We propose to manage uncertainties in autonomous, long-duration operations of multiagent systems via a modified Honeybee Foraging (HBF) behavioral scheme. The resulting system, Autonomous Persistent Intelligent Swarm (APIS),incorporates two new behaviors to ameliorate informational uncertainty. When “scouting”, agents are tasked based on informational quality and reliability rather than solely on priorities. When “dancing”, agents are tasked to rendezvous with other dancing agents to exchange information at close range, where successful communication is guaranteed. When coupled with uncertainty-aware modeling across agents, these behaviors improve situational awareness and resilience of the system, enabling it to function under more uncertain conditions arising during long-duration missions.

Autonomous systems↗

APIS: Honeybee Foraging Task Assignment for Use in Uncertain and Unreliable Environments

Multiagent Cyber-Physical-Human (CPH) systems in realistic environments operate under uncertain conditions. Communication among agents, aimed at reducing the uncertainty, is itself subject to uncertainty. We propose to manage uncertainties in autonomous, long-duration operations of multiagent systems via a modified Honeybee Foraging (HBF) behavioral scheme. The resulting system, Autonomous Persistent Intelligent Swarm (APIS),incorporates two new behaviors to ameliorate informational uncertainty. When “scouting”, agents are tasked based on informational quality and reliability rather than solely on priorities. When “dancing”, agents are tasked to rendezvous with other dancing agents to exchange information at close range, where successful communication is guaranteed. When coupled with uncertainty-aware modeling across agents, these behaviors improve situational awareness and resilience of the system, enabling it to function under more uncertain conditions arising during long-duration missions.

Autonomous systems↗