Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerant computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

NASA Tech Briefs, December 2005

Topics covered include: Video Mosaicking for Inspection of Gas Pipelines; Shuttle-Data-Tape XML Translator; Highly Reliable, High-Speed, Unidirectional Serial Data Links; Data-Analysis System for Entry, Descent, and Landing; Hybrid UV Imager Containing Face-Up AlGaN/GaN Photodiodes; Multiple Embedded Processors for Fault-Tolerant Computing; Hybrid Power Management; Magnetometer Based on Optoelectronic Microwave Oscillator; Program Predicts Time Courses of Human/ Computer Interactions; Chimera Grid Tools; Astronomer's Proposal Tool; Conservative Patch Algorithm and Mesh Sequencing for PAB3D; Fitting Nonlinear Curves by Use of Optimization Techniques; Tool for Viewing Faults Under Terrain; Automated Synthesis of Long Communication Delays for Testing; Solving Nonlinear Euler Equations With Arbitrary Accuracy; Self-Organizing-Map Program for Analyzing Multivariate Data; Tool for Sizing Analysis of the Advanced Life Support System; Control Software for a High-Performance Telerobot; Java Radar Analysis Tool; Architecture for Verifiable Software; Tool for Ranking Research Options; Enhanced, Partially Redundant Emergency Notification System; Close-Call Action Log Form; Task Description Language; Improved Small-Particle Powders for Plasma Spraying; Bonding-Compatible Corrosion Inhibitor for Rinsing Metals; Wipes, Coatings, and Patches for Detecting Hydrazines; Rotating Vessels for Growing Protein Crystals; Oscillating-Linear-Drive Vacuum Compressor for CO2; Mechanically Biased, Hinged Pairs of Piezoelectric Benders; Apparatus for Precise Indium-Bump Bonding of Microchips; Radiation Dosimetry via Automated Fluorescence Microscopy; Multistage Magnetic Separator of Cells and Proteins; Elastic-Tether Suits for Artificial Gravity and Exercise; Multichannel Brain-Signal-Amplifying and Digitizing System; Ester-Based Electrolytes for Low-Temperature Li-Ion Cells; Hygrometer for Detecting Water in Partially Enclosed Volumes; Radio-Frequency Plasma Cleaning of a Penning Malmberg Trap; Reduction of Flap Side Edge Noise - the Blowing Flap; and Preventing Accidental Ignition of Upper-Stage Rocket Motors.

Source record↗

Markov reward processes

Numerous applications in the area of computer system analysis can be effectively studied with Markov reward models. These models describe the behavior of the system with a continuous-time Markov chain, where a reward rate is associated with each state. In a reliability/availability model, upstates may have reward rate 1 and down states may have reward rate zero associated with them. In a queueing model, the number of jobs of certain type in a given state may be the reward rate attached to that state. In a combined model of performance and reliability, the reward rate of a state may be the computational capacity, or a related performance measure. Expected steady-state reward rate and expected instantaneous reward rate are clearly useful measures of the Markov reward model. More generally, the distribution of accumulated reward or time-averaged reward over a finite time interval may be determined from the solution of the Markov reward model. This information is of great practical significance in situations where the workload can be well characterized (deterministically, or by continuous functions e.g., distributions). The design process in the development of a computer system is an expensive and long term endeavor. For aerospace applications the reliability of the computer system is essential, as is the ability to complete critical workloads in a well defined real time interval. Consequently, effective modeling of such systems must take into account both performance and reliability. This fact motivates our use of Markov reward models to aid in the development and evaluation of fault tolerant computer systems.

Smith, R. M.↗

Current research activities at the NASA-sponsored Illinois Computing Laboratory of Aerospace Systems and Software

The Illinois Computing Laboratory of Aerospace Systems and Software (ICLASS) was established to: (1) pursue research in the areas of aerospace computing systems, software and applications of critical importance to NASA, and (2) to develop and maintain close contacts between researchers at ICLASS and at various NASA centers to stimulate interaction and cooperation, and facilitate technology transfer. Current ICLASS activities are in the areas of parallel architectures and algorithms, reliable and fault tolerant computing, real time systems, distributed systems, software engineering and artificial intelligence.

Smith, Kathryn A.↗

Three real-time architectures - A study using reward models

Numerous applications in the area of computer system analysis can be effectively studied with Markov reward models. These models describe the evolutionary behavior of the computer system by a continuous-time Markov chain, and a reward rate is associated with each state. In reliability/availability models, upstates have reward rate 1, and down states have reward rate zero associated with them. In a combined model of performance and reliability, the reward rate of a state may be the computational capacity, or a related performance measure. Steady-state expected reward rate and expected instantaneous reward rate are clearly useful measures which can be extracted from the Markov reward model. The diversity of areas where Markov reward models may be used is illustrated with a comparative study of three examples of interest to the fault tolerant computing community.

Sjogren, J. A.↗

Space station Ada runtime support for nested atomic transactions

The Space Station Data Management System (DMS), associated computing subsystems, and applications have varying degrees of reliability associated with their operation. A model has been developed (McKay '86) which allows the DMS runtime environment to appear as an Ada virtual machine to applications executing within it. This model is modular, flexible, and dynamically configurable to allow for evolution and growth over time. Support for Fault-tolerant computing is included within this model. The basic primitive involved in this support is based on atomic actions (Grey '78). An atomic action possesses two fundamental properties: (1) it is indivisible with respect to concurrent actions, and (2) it is indivisible with respect to failure. A transaction is a collection of atomic actions which collectively appear to be one action. Transactions may be nested, providing even more powerful support for reliability. A proposed approach is described for providing support for nested atomic transactions within the Ada runtime model developed for the Space Station environment. The level of support is modular, flexible and dynamically configurable just like the overall runtime support environment.

Monteiro, Edward J.↗

Validation of the SURE Program, phase 1

Presented are the results of the first phase in the validation of the SURE (Semi-Markov Unreliability Range Evaluator) program. The SURE program gives lower and upper bounds on the death-state probabilities of a semi-Markov model. With these bounds, the reliability of a semi-Markov model of a fault-tolerant computer system can be analyzed. For the first phase in the validation, fifteen semi-Markov models were solved analytically for the exact death-state probabilities and these solutions compared to the corresponding bounds given by SURE. In every case, the SURE bounds covered the exact solution. The bounds, however, had a tendency to separate in cases where the recovery rate was slow or the fault arrival rate was fast.

Dotson, Kelly J.↗

Development and evaluation of a Fault-Tolerant Multiprocessor (FTMP) computer. Volume 3: FTMP test and evaluation

The experimental test and evaluation of the Fault-Tolerant Multiprocessor (FTMP) is described. Major objectives of this exercise include expanding validation envelope, building confidence in the system, revealing any weaknesses in the architectural concepts and in their execution in hardware and software, and in general, stressing the hardware and software. To this end, pin-level faults were injected into one LRU of the FTMP and the FTMP response was measured in terms of fault detection, isolation, and recovery times. A total of 21,055 stuck-at-0, stuck-at-1 and invert-signal faults were injected in the CPU, memory, bus interface circuits, Bus Guardian Units, and voters and error latches. Of these, 17,418 were detected. At least 80 percent of undetected faults are estimated to be on unused pins. The multiprocessor identified all detected faults correctly and recovered successfully in each case. Total recovery time for all faults averaged a little over one second. This can be reduced to half a second by including appropriate self-tests.

Lala, J. H.↗

A methodology for validating software reliability

A significant problem associated with fault tolerant computer system design is how to insure that there are no embedded software errors, so that an avionics computer system meets the required reliability level. To accomplish this, it is necessary to associate a 'probability of failure' with the operational flight program. It would be more correct to say that the probability of excitation of existing latent design errors within the program is required. In this sense, latent software errors are like latent hardware faults, and techniques that were previously used to measure the probability of failure of hardware due to fault latency can be used to measure the probability of failure of the software. A methodology was developed and applied to a flight control program that was known to operate in a well defined environment. The results indicated that the technique could be used to provide a final validation of the software to a specified reliability level and to evaluate the role of flight test in software validation.

Swern, Frederic L.↗

Formal specification and mechanical verification of SIFT - A fault-tolerant flight control system

The paper describes the methodology being employed to demonstrate rigorously that the SIFT (software-implemented fault-tolerant) computer meets its requirements. The methodology uses a hierarchy of design specifications, expressed in the mathematical domain of multisorted first-order predicate calculus. The most abstract of these, from which almost all details of mechanization have been removed, represents the requirements on the system for reliability and intended functionality. Successive specifications in the hierarchy add design and implementation detail until the PASCAL programs implementing the SIFT executive are reached. A formal proof that a SIFT system in a 'safe' state operates correctly despite the presence of arbitrary faults has been completed all the way from the most abstract specifications to the PASCAL program.

Melliar-Smith, P. M.↗

Fault-tolerance experiments with the JPL STAR computer.

Results of fault-tolerance experiments performed using an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a 'program rollback' provision to eliminate transient-caused errors. After a brief review of the specification of fault-tolerance with respect to transient faults, including a description of the method of injection of transient faults in software and system tests, fault-tolerance experiments carried out with this computer with regard to the determination of fault classes, software verification, system verification, and recovery stability are summarized. A test and repair processor is described which constitutes a special monitor unit of the computer and is used to obtain information for fault detection in the other subsystems of the computer and to ensure that proper recovery occurs when a fault is detected.

Avizienis, A.↗

Logic design for dynamic and interactive recovery.

Recovery in a fault-tolerant computer means the continuation of system operation with data integrity after an error occurs. This paper delineates two parallel concepts embodied in the hardware and software functions required for recovery; detection, diagnosis, and reconfiguration for hardware, data integrity, checkpointing, and restart for the software. The hardware relies on the recovery variable set, checking circuits, and diagnostics, and the software relies on the recovery information set, audit, and reconstruct routines, to characterize the system state and assist in recovery when required. Of particular utility is a handware unit, the recovery control unit, which serves as an interface between error detection and software recovery programs in the supervisor and provides dynamic interactive recovery.

Carter, W. C.↗

Management and design of long-life systems; Proceedings of the Symposium, Denver, Colo., April 24-26, 1973

The long life of Pioneer interplanetary spacecraft is considered along with a general accelerated methodology for long-life mechanical components, dependable long-lived household appliances, and the design and development philosophy to achieve reliability and long life in large turbine generators. Other topics discussed include an integrated management approach to long life in space, artificial heart reliability factors, and architectural concepts and redundancy techniques in fault-tolerant computers. Individual items are announced in this issue.

Schurmeier, H. M.↗

Tug avionics system overview

The recently defined Tug avionics system takes maximum advantage of projected technology advances to attain: low system weight; power system capacity essentially independent of mission duration; sensors for rendezvous, docking, and navigation update; all attitude communications; onboard checkout and redundancy management; and modular fault tolerant computer control. The requirements, selection trades, and configurations are discussed for the major subsystems: data management, guidance navigation and control, communications, rendezvous and docking, and electrical power. The integrated avionics system and the interfaces with the payload, Shuttle and ground are described.

Raaberg, M. T.↗

Formal methods for achieving reliable software

Requirements for reliable avionic systems are discussed in terms of the effectiveness of programming methodology. The need for methods to cope with the complexity of critical real-time systems is emphasized. Some general concepts about formal methods are presented and an example is given of the SRI hierarchical development methodology taken from the executive system of the SIFT fault tolerant computer. Formal methods with alternatives are compared and the prospects for introducing formal methods into practice are considered.

Goldberg, J.↗

Self-Checking Memory Interface

Memory-interface integrated circuit not only detects errors in data from other circuits but also detects errors within itself. Memory-interface chip encodes 16-bit words with Hamming code for single-error correction or double-error detection. Chip used in fault-tolerant computers under development by NASA.

Sievers, M. W.↗

Data management

The following tasks were prioritized: software acquisition management plan; space station flight data system architectural study; space station user data system interface; automation of software development process; automation of software testing; distributed data base management; ADA (automated data acquisition) evaluation and transition and planning; network operating system software; fault tolerant computer validation methodology for onboard data management system; systems integration; artificial intelligence/expert systems; space station data network concept; space station standard interface protocols; space station data networks systems; integrated software development facility; and language trade studies.

Love, G.↗

Validation of a fault-tolerant clock synchronization system

A validation method for the synchronization subsystem of a fault tolerant computer system is investigated. The method combines formal design verification with experimental testing. The design proof reduces the correctness of the clock synchronization system to the correctness of a set of axioms which are experimentally validated. Since the reliability requirements are often extreme, requiring the estimation of extremely large quantiles, an asymptotic approach to estimation in the tail of a distribution is employed.

Butler, R. W.↗