Search NASA⌕ Search

SEARCH · Search NASA

Results for “hardware reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Highly Reliable, High-Speed, Unidirectional Serial Data Links

Highly reliable, high-speed, unidirectional serial data-communication subsystems have been proposed to be installed in an upgrade of the computing systems aboard the space shuttles. The basic design concept of these serial data links is also adaptable to terrestrial use in applications in which there are requirements for highly reliable serial data communications. The hardware and software aspects of the architecture of the data links are dictated largely by a requirement, in the original space-shuttle application, for one computer to monitor the memory transactions and memory contents of other computers in real time with high reliability and without reliance on requests for retransmission. To minimize weight while affording a capability to transfer data at a required rate of 2.56 x 10(exp 8) bits per second, it was decided that the links would be serial ones of the fiber-channel type. [Fiber channel denotes a type of serial computer bus that is used to connect a computer (usually a supercomputer) with a high-speed data storage device. Depending on the specific application, the physical connection between the transmitter and receiver could be made via an optical fiber or a twisted pair of wires.] Heretofore, fiber-channel links have ordinarily been bidirectional and have operated under protocols that provide for receiving stations to detect errors and request retransmission when necessary. In the present case, the time taken by processing to request retransmission would conflict with the requirement for real-time transfer of data. To ensure reliability without retransmission, a link according to the proposal would utilize a modified version of the normal fiberchannel character set in conjunction with forward error correction by means of a Reed-Solomon code (see figure). The Reed-Solomon encoding and decoding and the translations between the normal and modified character sets would be effected by logic circuitry external to the fiber-channel transmitter and receiver, which would be commercial off-the-shelf units. The receiving end of the link could detect and correct errors at a rate as high as 4 million times per second, if necessary. The receiver detects uncorrectable double-byte errors. It has been estimated that uncorrectable-error rate would amount to one failure in about 10(exp 19) characters.

Cole, Robert M.↗

On the feasibility of a spaceborne fault-tolerant hypercube

The feasibility of implementing a fault-tolerant hypercube architecture for space applications is discussed. Node-level architectures and designs are considered and a first-order reliability model is presented. It is shown how error recovery can be implemented using program rollback or roll-forward techniques. Shared memory augmentations to the message-passing structure can be used to get around the inefficiencies of multicomputers to provide efficient use of hardware to achieve the needed reliabilities while maintaining performance.

Rennels, David A.↗

Reliability and Probabilistic Risk Assessment - How They Play Together

PRA methodology is one of the probabilistic analysis methods that NASA brought from the nuclear industry to assess the risk of LOM, LOV and LOC for launch vehicles. PRA is a system scenario based risk assessment that uses a combination of fault trees, event trees, event sequence diagrams, and probability and statistical data to analyze the risk of a system, a process, or an activity. It is a process designed to answer three basic questions: What can go wrong? How likely is it? What is the severity of the degradation? Since 1986, NASA, along with industry partners, has conducted a number of PRA studies to predict the overall launch vehicles risks. Planning Research Corporation conducted the first of these studies in 1988. In 1995, Science Applications International Corporation (SAIC) conducted a comprehensive PRA study. In July 1996, NASA conducted a two-year study (October 1996 - September 1998) to develop a model that provided the overall Space Shuttle risk and estimates of risk changes due to proposed Space Shuttle upgrades. After the Columbia accident, NASA conducted a PRA on the Shuttle External Tank (ET) foam. This study was the most focused and extensive risk assessment that NASA has conducted in recent years. It used a dynamic, physics-based, integrated system analysis approach to understand the integrated system risk due to ET foam loss in flight. Most recently, a PRA for Ares I launch vehicle has been performed in support of the Constellation program. Reliability, on the other hand, addresses the loss of functions. In a broader sense, reliability engineering is a discipline that involves the application of engineering principles to the design and processing of products, both hardware and software, for meeting product reliability requirements or goals. It is a very broad design-support discipline. It has important interfaces with many other engineering disciplines. Reliability as a figure of merit (i.e. the metric) is the probability that an item will perform its intended function(s) for a specified mission profile. In general, the reliability metric can be calculated through the analyses using reliability demonstration and reliability prediction methodologies. Reliability analysis is very critical for understanding component failure mechanisms and in identifying reliability critical design and process drivers. The following sections discuss the PRA process and reliability engineering in detail and provide an application where reliability analysis and PRA were jointly used in a complementary manner to support a Space Shuttle flight risk assessment.

Safie, Fayssal M.↗

Failure modes, effects and criticality analyses.

Failure mode, effects and criticality analyses were developed by NASA as a means of assuring that hardware built for space applications has the desired reliability characteristics. The failure mode and effects analysis is a qualitative reliability technique for systematically analyzing each possible failure mode within a hardware system, and identifying the resulting effect on that system, the mission and personnel. The criticality analysis is a quantitative procedure which ranks the critical failure modes according to their probability of occurrence. This paper describes the failure modes, effects analysis and the criticality analysis. It employs a simple hardware system, not related to the aerospace field, to illustrate the method. It encourages application of this type of analysis to industrial development programs outside the aerospace and defense complex.

Jordan, W. E.↗

Space vehicle onboard command encoder

A flexible onboard encoder system was designed for the space shuttle. The following areas were covered: (1) implementation of the encoder design into hardware to demonstrate the various encoding algorithms/code formats, (2) modulation techniques in a single hardware package to maintain comparable reliability and link integrity of the existing link systems and to integrate the various techniques into a single design using current technology. The primary function of the command encoder is to accept input commands, generated either locally onboard the space shuttle or remotely from the ground, format and encode the commands in accordance with the payload input requirements and appropriately modulate a subcarrier for transmission by the baseband RF modulator. The following information was provided: command encoder system design, brassboard hardware design, test set hardware and system packaging, and software.

Source record↗

Issues in designing transport layer multicast facilities

Multicasting denotes a facility in a communications system for providing efficient delivery from a message's source to some well-defined set of locations using a single logical address. While modem network hardware supports multidestination delivery, first generation Transport Layer protocols (e.g., the DoD Transmission Control Protocol (TCP) (15) and ISO TP-4 (41)) did not anticipate the changes over the past decade in underlying network hardware, transmission speeds, and communication patterns that have enabled and driven the interest in reliable multicast. Much recent research has focused on integrating the underlying hardware multicast capability with the reliable services of Transport Layer protocols. Here, we explore the communication issues surrounding the design of such a reliable multicast mechanism. Approaches and solutions from the literature are discussed, and four experimental Transport Layer protocols that incorporate reliable multicast are examined.

Dempsey, Bert J.↗

Instrumentation for controlling and monitoring environmental control and life support systems

Advanced Instrumentation concepts for improving performance of manned spacecraft Environmental Control and Life Support Systems (EC/LSS) have been developed at Life Systems, Inc. The difference in specific EC/LSS instrumentation requirements and hardware during the transition from exploratory development to flight production stages are discussed. Details of prior control and monitor instrumentation designs are reviewed and an advanced design presented. The latter features a minicomputer-based approach having the flexibility to meet process hardware test programs and the capability to be refined to include the control dynamics and fault diagnostics needed in future flight systems where long duration, reliable operation requires in-flight hardware maintenance. The emphasis is on lower EC/LSS hardware life cycle costs by simplicity in instrumentation and using it to save crew time during flight operation.

Yang, P. Y.↗

The role of reliability graph models in assuring dependable operation of complex hardware/software systems

The complexity of computer systems currently being designed for critical applications in the scientific, commercial, and military arenas requires the development of new techniques for utilizing models of system behavior in order to assure 'ultra-dependability'. The complexity of these systems, such as Space Station Freedom and the Air Traffic Control System, stems from their highly integrated designs containing both hardware and software as critical components. Reliability graph models, such as fault trees and digraphs, are used frequently to model hardware systems. Their applicability for software systems has also been demonstrated for software safety analysis and the analysis of software fault tolerance. This paper discusses further uses of graph models in the design and implementation of fault management systems for safety critical applications.

Patterson-Hine, F. A.↗

The implementation and use of Ada on distributed systems with high reliability requirements

The use and implementation of Ada in distributed environments in which the hardware components are assumed to be unreliable is investigated. The possibility that a distributed system can be programmed entirely in Ada so that the individual tasks of the system are unconcerned with which processor they are executing on, and that failures can occur in the underlying hardware is considered. The reduced cost of computer hardware and the advantages of distributed processing (for example, increased reliability through redundancy and greater flexibility) indicate that many aerospace computer systems can be distributed. The use of Ada and distributed systems is a good combination for aerospace embedded systems.

Knight, J. C.↗

The J-2X Upper Stage Engine: From Heritage to Hardware

NASA's Global Exploration Strategy requires safe, reliable, robust, efficient transportation to support sustainable operations from Earth to orbit and into the far reaches of the solar system. NASA selected the Ares I crew launch vehicle and the Ares V cargo launch vehicle to provide that transportation. Guiding principles in creating the architecture represented by the Ares vehicles were the maximum use of heritage hardware and legacy knowledge, particularly Space Shuttle assets, and commonality between the Ares vehicles where possible to streamline the hardware development approach and reduce programmatic, technical, and budget risks. The J-2X exemplifies those goals. It was selected by the Exploration Systems Architecture Study (ESAS) as the upper stage propulsion for the Ares I Upper Stage and the Ares V Earth Departure Stage (EDS). The J-2X is an evolved version ofthe historic J-2 engine that successfully powered the second stage of the Saturn I launch vehicle and the second and third stages of the Saturn V launch vehicle. The Constellation architecture, however, requires performance greater than its predecessor. The new architecture calls for larger payloads delivered to the Moon and demands greater loss of mission reliability and numerous other requirements associated with human rating that were not applied to the original J-2. As a result, the J-2X must operate at much higher temperatures, pressures, and flow rates than the heritage J-2, making it one of the highest performing gas generator cycle engines ever built, approaching the efficiency of more complex stage combustion engines. Development is focused on early risk mitigation, component and subassembly test, and engine system test. The development plans include testing engine components, including the subscale injector, main igniter, powerpack assembly (turbopumps, gas generator and associated ducting and structural mounts), full-scale gas generator, valves, and control software with hardware-in-the-loop. Testing expanded in 2007, accompanied by the refinement of the design through several key milestones. This paper discusses those 2007 tests and milestones, as well as updates key developments in 2008.

Byrd, THomas↗

System applications of the fault tolerant memory

Conventional memory technologies currently employed in aerospace applications contribute at least fifty percent to system unreliability (where the system includes CPU, I/O and memory). A fault tolerant memory performs both error correction and memory replacement at the bit plane level. To determine the effects of system design of using a fault tolerant memory in space applications, analysis was performed to determine tradeable hardware configurations that meet the reliability goals of each program. The candidate configurations, which included redundant elements of the computer system with both conventional and fault tolerant memories, were then traded in terms of selection criteria of cost, weight, volume, and power. These trade studies demonstrated that a fault tolerant memory provided significant advantages in terms of cost, weight, and volume. The memory selected for this analysis was a recently developed five fault tolerant memory.

Murphy, L. J.↗

Performance characterization of a Bosch CO sub 2 reduction subsystem

The performance of Bosch hardware at the subsystem level (up to five-person capacity) in terms of five operating parameters was investigated. The five parameters were: (1) reactor temperature, (2) recycle loop mass flow rate, (3) recycle loop gas composition (percent hydrogen), (4) recycle loop dew point and (5) catalyst density. Experiments were designed and conducted in which the five operating parameters were varied and Bosch performance recorded. A total of 12 carbon collection cartridges provided over approximately 250 hours of operating time. Generally, one cartridge was used for each parameter that was varied. The Bosch hardware was found to perform reliably and reproducibly. No startup, reaction initiation or carbon containment problems were observed. Optimum performance points/ranges were identified for the five parameters investigated. The performance curves agreed with theoretical projections.

Heppner, D. B.↗

Design of a modular digital computer system

A Central Control Element (CCE) module which controls the Automatically Reconfigurable Modular System (ARMS) and allows both redundant processing and multi-computing in the same computer with real time mode switching, is discussed. The same hardware is used for either reliability enhancement, speed enhancement, or for a combination of both.

Source record↗

The fault-tolerant multiprocessor computer

The development and evaluation of fault-tolerant computer architectures and software-implemented fault tolerance (SIFT) for use in advanced NASA vehicles and potentially in flight-control systems are described in a collection of previously published reports prepared for NASA. Topics addressed include the principles of fault-tolerant multiprocessor (FTMP) operation; processor and slave regional designs; FTMP executive, facilities, acceptance-test/diagnostic, applications, and support software; FTM reliability and availability models; SIFT hardware design; and SIFT validation and verification.

Smith, T. B., III↗

The JPL/KSC telerobotic inspection demonstration

An ASEA IRB90 robotic manipulator with attached inspection cameras was moved through a Space Shuttle Payload Assist Module (PAM) Cradle under computer control. The Operator and Operator Control Station, including graphics simulation, gross-motion spatial planning, and machine vision processing, were located at JPL. The Safety and Support personnel, PAM Cradle, IRB90, and image acquisition system, were stationed at the Kennedy Space Center (KSC). Images captured at KSC were used both for processing by a machine vision system at JPL, and for inspection by the JPL Operator. The system found collision-free paths through the PAM Cradle, demonstrated accurate knowledge of the location of both objects of interest and obstacles, and operated with a communication delay of two seconds. Safe operation of the IRB90 near Shuttle flight hardware was obtained both through the use of a gross-motion spatial planner developed at JPL using artificial intelligence techniques, and infrared beams and pressure sensitive strips mounted to the critical surfaces of the flight hardward at KSC. The Demonstration showed that telerobotics is effective for real tasks, safe for personnel and hardware, and highly productive and reliable for Shuttle payload operations and Space Station external operations.

Mittman, David↗

Study on advanced information processing system

Issues related to the reliability of a redundant system with large main memory are addressed. In particular, the Fault-Tolerant Processor (FTP) for Advanced Launch System (ALS) is used as a basis for our presentation. When the system is free of latent faults, the probability of system crash due to nearly-coincident channel faults is shown to be insignificant even when the outputs of computing channels are infrequently voted on. In particular, using channel error maskers (CEMs) is shown to improve reliability more effectively than increasing the number of channels for applications with long mission times. Even without using a voter, most memory errors can be immediately corrected by CEMs implemented with conventional coding techniques. In addition to their ability to enhance system reliability, CEMs--with a low hardware overhead--can be used to reduce not only the need of memory realignment, but also the time required to realign channel memories in case, albeit rare, such a need arises. Using CEMs, we have developed two schemes, called Scheme 1 and Scheme 2, to solve the memory realignment problem. In both schemes, most errors are corrected by CEMs, and the remaining errors are masked by a voter.

Shin, Kang G.↗

Autonomous berthing/unberthing of a Work Attachment Mechanism/Work Attachment Fixture (WAM/WAF)

Discussed here is the autonomous berthing of a Work Attachment Mechanism/Work Attachment Fixture (WAM/WAF) developed by NASA for berthing and docking applications in space. The WAM/WAF system enables fast and reliable berthing (unberthing) of space hardware. A successful operation of the WAM/WAF requires that the WAM motor velocity be precisely controlled. The operating principle and the design of the WAM/WAF is described as well as the development of a control system used to regulate the WAM motor velocity. The results of an experiment in which the WAM/WAF is used to handle an orbital replacement unit are given.

Nguyen, Charles C.↗