Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerant computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34

Performance measures for multiprocessor controllers

Performance measures to characterize fault tolerant multiprocessors used in the control of critical processes are considered. Our performance indices are based on controller response time. By relating this to the needs of the application, we have been able to derive indices that faithfully reflect the performance of the multiprocessor in the context of the application, that permit the objective comparison of rival computer systems, and that can either be definitively estimated or objectively measured. An example of a controller in an idealized satellite application is provided.

Krishna, C. M.↗

Optimal maintenance center inventories for fault-tolerant repairable systems

A probabilistic approach is taken to determine the optimal repairable parts inventory for a maintenance center, servicing machines which contain several m-out-of-n systems of different parts, with a constraint on the total inventory investment. A model, based on the discrete Markov process, accounts for a typical ultrareliable avionics system, such as one presently being developed by NASA. The dynamic programming algorithm for minimizing the stockout and holding costs is applied to an exemplary maintenance center, and solutions for single-item and multi-item cases are given. The computational burden is noted to be reasonable and a computer program is used to generate optimal solutions.

Lawrence, S. H.↗

A Poisson process approximation for generalized K-5 confidence regions

One-sided confidence regions for continuous cumulative distribution functions are constructed using empirical cumulative distribution functions and the generalized Kolmogorov-Smirnov distance. The band width of such regions becomes narrower in the right or left tail of the distribution. To avoid tedious computation of confidence levels and critical values, an approximation based on the Poisson process is introduced. This aproximation provides a conservative confidence region; moreover, the approximation error decreases monotonically to 0 as sample size increases. Critical values necessary for implementation are given. Applications are made to the areas of risk analysis, investment modeling, reliability assessment, and analysis of fault tolerant systems.

Arsham, H.↗

Real-time closed-loop simulation and upset evaluation of control systems in harsh electromagnetic environments

Digital control systems for applications such as aircraft avionics and multibody systems must maintain adequate control integrity in adverse as well as nominal operating conditions. For example, control systems for advanced aircraft, and especially those with relaxed static stability, will be critical to flight and will, therefore, have very high reliability specifications which must be met regardless of operating conditions. In addition, multibody systems such as robotic manipulators performing critical functions must have control systems capable of robust performance in any operating environment in order to complete the assigned task reliably. Severe operating conditions for electronic control systems can result from electromagnetic disturbances caused by lightning, high energy radio frequency (HERF) transmitters, and nuclear electromagnetic pulses (NEMP). For this reason, techniques must be developed to evaluate the integrity of the control system in adverse operating environments. The most difficult and illusive perturbations to computer-based control systems that can be caused by an electromagnetic environment (EME) are functional error modes that involve no component damage. These error modes are collectively known as upset, can occur simultaneously in all of the channels of a redundant control system, and are software dependent. Upset studies performed to date have not addressed the assessment of fault tolerant systems and do not involve the evaluation of a control system operating in a closed-loop with the plant. A methodology for performing a real-time simulation of the closed-loop dynamics of a fault tolerant control system with a simulated plant operating in an electromagnetically harsh environment is presented. In particular, considerations for performing upset tests on the controller are discussed. Some of these considerations are the generation and coupling of analog signals representative of electromagnetic disturbances to a control system under test, analog data acquisition, and digital data acquisition from fault tolerant systems. In addition, a case study of an upset test methodology for a fault tolerant electromagnetic aircraft engine control system is presented.

Belcastro, Celeste M.↗

Dynamical Decoupling of Crosstalk on Superconducting Qubit Devices

Current NISQ devices are prone to errors. In order to be used for practical applications or achieve fault-tolerant thresholds, strategies to suppress error rates will be needed to maximize the potential of noisy devices. Dynamical decoupling (DD) is one such strategy for suppressing — or at least alleviating — the effects of decoherence, in which sequences of pulses are applied to qubits to decouple their interaction with the environment. Through experimental runs performed on several Rigetti quantum computing units (QPUs), we first demonstrate that DD is capable of improving coherence times for isolated qubits, as well as suppressing errors caused by the ZZ coupling between pairs of qubits. Extending this framework to cycles containing2-qubit gates, we show that DD can be inserted to decouple qubits from crosstalk occurring during neighboring 2-qubit gates, and demonstrate the efficacy of this procedure on quantum approximate optimization algorithm (QAOA) circuits. We also explore the usage of tailored DD sequences for the suppression of characterized error channels. We are grateful for support from the NASA Ames Research Center and from the DARPA ONISQ program under interagency agreement IAA 8839,Annex 114. HYH is supported by the USRA Feynman QuantumAcademy funded by the NAMS R&D Student Program and a UCHellman Fellowship. JS, ZGI and ZW are supported by USRA NASAAcademic Mission Service (NNA16BD14C).

Dynamical decoupling↗

Fault tolerance in onboard processors - Protecting efficient FDM demultiplexers

The application of convolutional codes to protect demultiplexer filter banks is demonstrated analytically for efficient implementations. An overview is given of the parameters for the efficient implementations of filter banks, and real convolutional codes are discussed in terms of DSP operations. Methods for composite filtering and parity generation are outlined, and attention is given to the protection of polyphase filter demultiplexing systems. Real convolutional codes can be applied to protect demultiplexer filter banks by employing two forms of low-rate parity calculation to each filter bank. The parity values are computed either by the output with an FIR parity filter or in parallel with the normal processing by a composite filter. Hardware similarities between the filter bank and the main demultiplexer bank permit efficient redeployment of the processing resources to the main processing function in any configuration.

Redinbo, Robert↗

Fault-tolerance memory system architecture for radiation induced errors

A fault-tolerant memory (FTM) architecture is presented which can be used to overcome soft memory errors induced by alpha particles, cosmic radiation, or other random sources. The characteristics of the FTM are presented, a mathematical model is developed, and numerical examples are considered to illustrate the effectiveness of the approach. The FTM architecture has been incorporated in the NASA Standard Spacecraft Computer (NSSC-II) which will be employed in a variety of future space payloads and experiments.

White, J. B., Jr.↗

Networked Array Recorder (NeAR) Microphones for Field-Deployed Phased Arrays

An innovative edge-computing concept known as NeAR (Networked Array Recorder) has been developed to provide enhancements to existing field-deployable microphone phased arrays utilized for aeroacoustic flyover measurements of airframe and propulsive noise sources. The proposed system allows for the elimination of multiple miles of sensor wiring in an array installation, thereby improving the scalability of the overall system, increasing the fault-tolerance of the hardware, and reducing the effort needed to build-up and tear-down an array in the field. A demonstration of the NeAR concept was performed at Edwards Air Force Base in California in March – April, 2018, where twelve individual NeAR microphones were deployed as a piggyback on a conventional phased array system deployed for airframe noise flyover testing. The microphones operated successfully during the demonstration with good time history and spectral correlations shown between the NeAR units and conventional microphones located nearby in the array. The NeAR concept has spinoffs beyond its use for phased arrays, including applications in remote environmental sensing and noise monitoring.

Cull.iton, William G.↗

Single-ancilla ground state preparation via Lindbladians

We design a quantum algorithm for ground state preparation in the early fault tolerant regime. As a Monte Carlo style quantum algorithm, our method features a Lindbladian where the target state is stationary. The construction of this Lindbladian is algorithmic and should not be seen as a specific approximation to some weakly coupled system-bath dynamics in nature. Our algorithm can be implemented using just one ancilla qubit and efficiently simulated on a quantum computer. It can prepare the ground state even when the initial state has zero overlap with the ground state, bypassing the most significant limitation of methods like quantum phase estimation. As a variant, we also propose a discrete-time algorithm, demonstrating even better efficiency and providing a near-optimal simulation cost depending on the desired evolution time and precision. Numerical simulations using Ising and Hubbard models demonstrate the efficacy and applicability of our method. Published by the American Physical Society 2024

Ding, Zhiyan (ORCID:000000018863403X)↗

SIRU development. Volume 3: Software description and program documentation

The development and initial evaluation of a strapdown inertial reference unit (SIRU) system are discussed. The SIRU configuration is a modular inertial subsystem with hardware and software features that achieve fault tolerant operational capabilities. The SIRU redundant hardware design is formulated about a six gyro and six accelerometer instrument module package. The six axes array provides redundant independent sensing and the symmetry enables the formulation of an optimal software redundant data processing structure with self-contained fault detection and isolation (FDI) capabilities. The basic SIRU software coding system used in the DDP-516 computer is documented.

Oehrle, J.↗

Human factors aspects of control room design

A plan for the design and analysis of a multistation control room is reviewed. It is found that acceptance of the computer based information system by the uses in the control room is mandatory for mission and system success. Criteria to improve computer/user interface include: match of system input/output with user; reliability, compatibility and maintainability; easy to learn and little training needed; self descriptive system; system under user control; transparent language, format and organization; corresponds to user expectations; adaptable to user experience level; fault tolerant; dialog capability user communications needs reflected in flexibility, complexity, power and information load; integrated system; and documentation.

Jenkins, J. P.↗

Hyperswitch communication network

The Hyperswitch Communication Network (HCN) is a large scale parallel computer prototype being developed at JPL. Commercial versions of the HCN computer are planned. The HCN computer being designed is a message passing multiple instruction multiple data (MIMD) computer, and offers many advantages in price-performance ratio, reliability and availability, and manufacturing over traditional uniprocessors and bus based multiprocessors. The design of the HCN operating system is a uniquely flexible environment that combines both parallel processing and distributed processing. This programming paradigm can achieve a balance among the following competing factors: performance in processing and communications, user friendliness, and fault tolerance. The prototype is being designed to accommodate a maximum of 64 state of the art microprocessors. The HCN is classified as a distributed supercomputer. The HCN system is described, and the performance/cost analysis and other competing factors within the system design are reviewed.

Peterson, J.↗

Autonomous fault-tolerant attitude reference system using DTGs in symmetrically skewed configuration

A novel symmetrically skewed configuration for an attitude reference system (ARS) using three dynamically tuned gyros (DTGs) is developed. Simple schemes for autonomous detection and identification of a faulty DTG in real time and subsequent reconfiguration of the attitude estimation algorithm are proposed. The performance of the present configuration is shown to be better than that of configurations proposed earlier, and it is shown to have better features. It tolerates all types of failures of DTG failures, requires very simple computations, and gives less error in attitude estimate than the other configurations.

Murugesan, S.↗

Reduction Of Sizes Of Semi-Markov Reliability Models

Trimming technique reduces computational effort by order of magnitude while introducing negligible error. Error bound depends on only three parameters from semi-Markov model: maximum sum of rates for failure transitions leaving any state, maximum average holding time for recovery-mode state, and operating time for system. Error bound computed before any model generated, enabling modeler to decide immediately whether or not model can be trimmed. Trimming procedure specified by precise and easy description, making it easy to include trimming procedure in program generating mathematical models for use in assessing reliability. Typical application of technique in design of digital control systems required to be extremely reliable. In addition to aerospace applications, fault-tolerant design has growing importance in wide range of industrial applications.

White, Allan L.↗

FINDS: A fault inferring nonlinear detection system. User's guide

The computer program FINDS is written in FORTRAN-77, and is intended for operation on a VAX 11-780 or 11-750 super minicomputer, using the VMS operating system. The program detects, isolates, and compensates for failures in navigation aid instruments and onboard flight control and navigation sensors of a Terminal Configured Vehicle aircraft in a Microwave Landing System environment. In addition, FINDS provides sensor fault tolerant estimates for the aircraft states which are then used by an automatic guidance and control system to land the aircraft along a prescribed path. FINDS monitors for failures by evaluating all sensor outputs simultaneously using the nonlinear analytic relationships between the various sensor outputs arising from the aircraft point mass equations of motion. Hence, FINDS is an integrated sensor failure detection and isolation system.

Lancraft, R. E.↗

Permutation codes for the state assignment of fault tolerant sequential machines

A new fault-tolerant state assignment method is suggested for synchronous sequential machines. It is assumed that the inputs are fault free and that for no input it is possible to reach all or most of the states, whose number may be fairly large. Error correcting codes for the state assignment are generated by permutations of a chosen linear code. A state assignment algorithm is developed and its computational complexity is estimated. Examples are given.

Chen, M.↗

Evaluation of fault-tolerant system performance by approximate techniques

An approximate method for calculating the statistics of the performance of a fault-tolerant system is developed. An approximate method is necessary because the statistical model of the system behavior is large-scale and the time horizon of interest encompasses many cycles of the Redundancy Management logic. In the development, a compact representation of the necessary information called the v-transform is introduced and discussed. Based upon this representation, an approximation that leads to a very efficient computational procedure is suggested and numerically analyzed. A very brief discussion of other related work is also presented.

Walker, B. K.↗