Search NASA⌕ Search

SEARCH · Search NASA

Results for “fault tolerant computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A methodology for validating software reliability

A significant problem associated with fault tolerant computer system design is how to insure that there are no embedded software errors, so that an avionics computer system meets the required reliability level. To accomplish this, it is necessary to associate a 'probability of failure' with the operational flight program. It would be more correct to say that the probability of excitation of existing latent design errors within the program is required. In this sense, latent software errors are like latent hardware faults, and techniques that were previously used to measure the probability of failure of hardware due to fault latency can be used to measure the probability of failure of the software. A methodology was developed and applied to a flight control program that was known to operate in a well defined environment. The results indicated that the technique could be used to provide a final validation of the software to a specified reliability level and to evaluate the role of flight test in software validation.

Swern, Frederic L.↗

Formal specification and mechanical verification of SIFT - A fault-tolerant flight control system

The paper describes the methodology being employed to demonstrate rigorously that the SIFT (software-implemented fault-tolerant) computer meets its requirements. The methodology uses a hierarchy of design specifications, expressed in the mathematical domain of multisorted first-order predicate calculus. The most abstract of these, from which almost all details of mechanization have been removed, represents the requirements on the system for reliability and intended functionality. Successive specifications in the hierarchy add design and implementation detail until the PASCAL programs implementing the SIFT executive are reached. A formal proof that a SIFT system in a 'safe' state operates correctly despite the presence of arbitrary faults has been completed all the way from the most abstract specifications to the PASCAL program.

Melliar-Smith, P. M.↗

Fault-tolerance experiments with the JPL STAR computer.

Results of fault-tolerance experiments performed using an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a 'program rollback' provision to eliminate transient-caused errors. After a brief review of the specification of fault-tolerance with respect to transient faults, including a description of the method of injection of transient faults in software and system tests, fault-tolerance experiments carried out with this computer with regard to the determination of fault classes, software verification, system verification, and recovery stability are summarized. A test and repair processor is described which constitutes a special monitor unit of the computer and is used to obtain information for fault detection in the other subsystems of the computer and to ensure that proper recovery occurs when a fault is detected.

Avizienis, A.↗

Logic design for dynamic and interactive recovery.

Recovery in a fault-tolerant computer means the continuation of system operation with data integrity after an error occurs. This paper delineates two parallel concepts embodied in the hardware and software functions required for recovery; detection, diagnosis, and reconfiguration for hardware, data integrity, checkpointing, and restart for the software. The hardware relies on the recovery variable set, checking circuits, and diagnostics, and the software relies on the recovery information set, audit, and reconstruct routines, to characterize the system state and assist in recovery when required. Of particular utility is a handware unit, the recovery control unit, which serves as an interface between error detection and software recovery programs in the supervisor and provides dynamic interactive recovery.

Carter, W. C.↗

Management and design of long-life systems; Proceedings of the Symposium, Denver, Colo., April 24-26, 1973

The long life of Pioneer interplanetary spacecraft is considered along with a general accelerated methodology for long-life mechanical components, dependable long-lived household appliances, and the design and development philosophy to achieve reliability and long life in large turbine generators. Other topics discussed include an integrated management approach to long life in space, artificial heart reliability factors, and architectural concepts and redundancy techniques in fault-tolerant computers. Individual items are announced in this issue.

Schurmeier, H. M.↗

Tug avionics system overview

The recently defined Tug avionics system takes maximum advantage of projected technology advances to attain: low system weight; power system capacity essentially independent of mission duration; sensors for rendezvous, docking, and navigation update; all attitude communications; onboard checkout and redundancy management; and modular fault tolerant computer control. The requirements, selection trades, and configurations are discussed for the major subsystems: data management, guidance navigation and control, communications, rendezvous and docking, and electrical power. The integrated avionics system and the interfaces with the payload, Shuttle and ground are described.

Raaberg, M. T.↗

Formal methods for achieving reliable software

Requirements for reliable avionic systems are discussed in terms of the effectiveness of programming methodology. The need for methods to cope with the complexity of critical real-time systems is emphasized. Some general concepts about formal methods are presented and an example is given of the SRI hierarchical development methodology taken from the executive system of the SIFT fault tolerant computer. Formal methods with alternatives are compared and the prospects for introducing formal methods into practice are considered.

Goldberg, J.↗

Self-Checking Memory Interface

Memory-interface integrated circuit not only detects errors in data from other circuits but also detects errors within itself. Memory-interface chip encodes 16-bit words with Hamming code for single-error correction or double-error detection. Chip used in fault-tolerant computers under development by NASA.

Sievers, M. W.↗

Data management

The following tasks were prioritized: software acquisition management plan; space station flight data system architectural study; space station user data system interface; automation of software development process; automation of software testing; distributed data base management; ADA (automated data acquisition) evaluation and transition and planning; network operating system software; fault tolerant computer validation methodology for onboard data management system; systems integration; artificial intelligence/expert systems; space station data network concept; space station standard interface protocols; space station data networks systems; integrated software development facility; and language trade studies.

Love, G.↗

Validation of a fault-tolerant clock synchronization system

A validation method for the synchronization subsystem of a fault tolerant computer system is investigated. The method combines formal design verification with experimental testing. The design proof reduces the correctness of the clock synchronization system to the correctness of a set of axioms which are experimentally validated. Since the reliability requirements are often extreme, requiring the estimation of extremely large quantiles, an asymptotic approach to estimation in the tail of a distribution is employed.

Butler, R. W.↗

Complementary-Logic Fault Detector

Circuit for checking two-line complementary-logic bits for single faults used as building block for self-checking memory interface for Hammingcoded data. Intended for such applications as fault-tolerant computing, data handling, and data transmission. Circuit performs exclusive-OR function. Many such circuits combined produce complete memory interface with both detection and correction abilities.

Wawrzynek, J. C.↗

Technologies for space station autonomy

This report presents an informal survey of experts in the field of spacecraft automation, with recommendations for which technologies should be given the greatest development attention for implementation on the initial 1990's NASA Space Station. The recommendations implemented an autonomy philosophy that was developed by the Concept Development Group's Autonomy Working Group during 1983. They were based on assessments of the technologies' likely maturity by 1987, and of their impact on recurring costs, non-recurring costs, and productivity. The three technology areas recommended for programmatic emphasis were: (1) artificial intelligence expert (knowledge based) systems and processors; (2) fault tolerant computing; and (3) high order (procedure oriented) computer languages. This report also describes other elements required for Station autonomy, including technologies for later implementation, system evolvability, and management attitudes and goals. The cost impact of various technologies is treated qualitatively, and some cases in which both the recurring and nonrecurring costs might be reduced while the crew productivity is increased, are also considered. Strong programmatic emphasis on life cycle cost and productivity is recommended.

Staehle, R. L.↗

Space station data system analysis/architecture study. Task 3: Trade studies, DR-5, volume 2

Results of a Space Station Data System Analysis/Architecture Study for the Goddard Space Flight Center are presented. This study, which emphasized a system engineering design for a complete, end-to-end data system, was divided into six tasks: (1); Functional requirements definition; (2) Options development; (3) Trade studies; (4) System definitions; (5) Program plan; and (6) Study maintenance. The Task inter-relationship and documentation flow are described. Information in volume 2 is devoted to Task 3: trade Studies. Trade Studies have been carried out in the following areas: (1) software development test and integration capability; (2) fault tolerant computing; (3) space qualified computers; (4) distributed data base management system; (5) system integration test and verification; (6) crew workstations; (7) mass storage; (8) command and resource management; and (9) space communications. Results are presented for each task.

Source record↗

Wavelength division multiplexing

Wavelength division multiplexing (WDM) represents an approach for expanding the communication capacity and for implementing special data techniques in a fiber optics system. This technology is implemented by adding optical sources of different wavelengths at optical transmitting locations. The present paper is concerned with some of the current efforts in WDM. WDM applications are related to long haul communications, local area data networks, spacecraft and aircraft data systems, fault tolerant computer networks, special sensor devices, high speed data processors, closed circuit and cable television, and submarine cable systems. Attention is given to the current state of wavelength division multiplexing applications, the availability and status of WDM components semiconductor lasers/transmitters, availability and status of fiber optic detectors/receivers, optical fibers/cables/connectors/taps/star couplers, wavelength multiplexers/demultiplexers, and future WDM for local area networks.

Hendricks, H. D.↗

Analysis of typical fault-tolerant architectures using HARP

Difficulties encountered in the modeling of fault-tolerant systems are discussed. The Hybrid Automated Reliability Predictor (HARP) approach to modeling fault-tolerant systems is described. The HARP is written in FORTRAN, consists of nearly 30,000 lines of codes and comments, and is based on behavioral decomposition. Using the behavioral decomposition, the dependability model is divided into fault-occurrence/repair and fault/error-handling models; the characteristics and combining of these two models are examined. Examples in which the HARP is applied to the modeling of some typical fault-tolerant systems, including a local-area network, two fault-tolerant computer systems, and a flight control system, are presented.

Bavuso, Salvatore J.↗

NAECON 87; Proceedings of the IEEE National Aerospace and Electronics Conference, Dayton, OH, May 18-22, 1987. Volumes 1, 2, 3, & 4

The present conference discusses topics in VLSI components and their packaging, signal processing, uses of cartographic data, data transmission, advanced avionics architectures, fiber-optics, information control and display, image processing, airborne radar and fire control, navigation, air data, Kalman filtering, power generation and control, spacecraft power structures, aircraft flying qualities, flight management, fault-tolerant computer architectures, actuation technologies, self-repairing flight control system technology, multivariable control, stability and control methods, and AFTI/F-16 flight test reports. Also discussed are the ADA/JOVIAL language and its applications, software acquisition and testing, advanced software concepts, software management, computer graphics and visual systems softwear, ADA in embedded avionics, 16- and 32-bit architectures, voice interaction applications, human/machine systems analysis, human factors and AI, mental workloads and displays, pilot acceleration protection research, communications system technology, space communications, reliability and maintainability, managerial techniques, engineering management, EM compatibility and nuclear hardening, expert systems, AI language/knowledge representation, expert system implementation, machine vision/optical processing, and advanced AI concepts and architectures.

Avionics↗

Transparent Ada rendezvous in a fault tolerant distributed system

There are many problems associated with distributing an Ada program over a loosely coupled communication network. Some of these problems involve the various aspects of the distributed rendezvous. The problems addressed involve supporting the delay statement in a selective call and supporting the else clause in a selective call. Most of these difficulties are compounded by the need for an efficient communication system. The difficulties are compounded even more by considering the possibility of hardware faults occurring while the program is running. With a hardware fault tolerant computer system, it is possible to design a distribution scheme and communication software which is efficient and allows Ada semantics to be preserved. An Ada design for the communications software of one such system will be presented, including a description of the services provided in the seven layers of an International Standards Organization (ISO) Open System Interconnect (OSI) model communications system. The system capabilities (hardware and software) that allow this communication system will also be described.

Racine, Roger↗