Search NASA⌕ Search

SEARCH · Search NASA

Results for “software failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Optimal integral controller with sensor failure accommodation

An Optimal Integral Controller that readily accommodates Sensor Failure - without resorting to (Kalman) filter or observer generation - has been designed. The system is based on Navy-sponsored research for the control of high performance aircraft. In conjunction with a NASA developed Numerical Optimization Code, the Integral Feedback Controller will provide optimal system response even in the case of incomplete state feedback. Hence, the need for costly replication of plant sensors is avoided since failure accommodation is effected by system software reconfiguration. The control design has been applied to a particularly ill-behaved, third-order system. Dominant-root design in the classical sense produced an almost 100 percent overshoot for the third-order system response. An application of the newly-developed Optimal Integral Controller - assuming all state information available - produces a response with no overshoot. A further application of the controller design - assuming a one-third sensor failure scenario - produced a slight overshoot response that still preserved the steady state time-point of the full-state feedback response. The control design should have wide application in space systems.

Alberts, T.↗

A Testbed for Evaluating Lunar Habitat Autonomy Architectures

A lunar outpost will involve a habitat with an integrated set of hardware and software that will maintain a safe environment for human activities. There is a desire for a paradigm shift whereby crew will be the primary mission operators, not ground controllers. There will also be significant periods when the outpost is uncrewed. This will require that significant automation software be resident in the habitat to maintain all system functions and respond to faults. JSC is developing a testbed to allow for early testing and evaluation of different autonomy architectures. This will allow evaluation of different software configurations in order to: 1) understand different operational concepts; 2) assess the impact of failures and perturbations on the system; and 3) mitigate software and hardware integration risks. The testbed will provide an environment in which habitat hardware simulations can interact with autonomous control software. Faults can be injected into the simulations and different mission scenarios can be scripted. The testbed allows for logging, replaying and re-initializing mission scenarios. An initial testbed configuration has been developed by combining an existing life support simulation and an existing simulation of the space station power distribution system. Results from this initial configuration will be presented along with suggested requirements and designs for the incremental development of a more sophisticated lunar habitat testbed.

Lawler, Dennis G.↗

Electrified Aircraft Propulsion Systems: Potential Failure Modes and Failure Mitigation Strategies

Electrified aircraft propulsion (EAP) systems hold great potential for the reduction of aircraft fuel burn, emissions, and noise. Currently, NASA and other organizations are actively working to identify and mature technologies necessary to bring EAP designs to reality. A requirement for the development of any civil aircraft and its systems is to ensure that potential hazards in the design are identified and appropriately mitigated to ensure that the system is safe. During aircraft development, a system safety assessment that consists of a functional hazard assessment is conducted to identify all potential failure conditions of each function, and classify those failures according to the severity of their effects on the aircraft or its occupants. The more severe a function's failure condition classification, the greater the development assurance level required for the function to ensure that the probability of the hazard is acceptably low. Today, aircraft engines and their control systems receive type certificate approval as a stand-alone system to signify their airworthiness. However, the complex coupling and distributed nature of EAP designs are expected to place added challenges on the certification of these systems. This presentation will provide an initial high-level review of the potential failure modes and hazards posed by a generic EAP system along with potential mitigation strategies for those failures. The EAP system is assumed to be a hybrid design consisting of gas turbine engines, mechanical drives, electric machines, power electronics and distribution systems, energy storage devices, and motor driven propulsors. The functionality provided by each of these EAP subsystems will be discussed along with the potential failure modes they may encounter. This will include a discussion of coupled failure effects, where a fault in one EAP subsystem effects the operation of other subsystems in the architecture. Next, potential failure mitigation strategies are discussed including both software-based and hardware-based mitigation strategies. The presentation will conclude with an example evaluation of the potential failure modes and mitigation strategies for a concept EAP system proposed by NASA.

Simon, Donald L.↗

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failures. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certification processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E., Jr.↗

Spacecraft Software Maintenance: An Effective Approach to Reducing Costs and Increasing Science Return

Flight software is a mission critical element of spacecraft functionality and performance. When ground operations personnel interface to a spacecraft, they are typically dealing almost entirely with the capabilities of onboard software. This software, even more than critical ground/flight communications systems, is expected to perform perfectly during all phases of spacecraft life. Due to the fact that it can be reprogrammed on-orbit to accommodate degradations or failures in flight hardware, new insights into spacecraft characteristics, new control options which permit enhanced science options, etc., the on- orbit flight software maintenance team is usually significantly responsible for the long term success of a science mission. Failure of flight software to perform as needed can result in very expensive operations work-around costs and lost science opportunities. There are three basic approaches to maintaining spacecraft software--namely using the original developers, using the mission operations personnel, or assembling a center of excellence for multi-spacecraft software maintenance. Not planning properly for flight software maintenance can lead to unnecessarily high on-orbit costs and/or unacceptably long delays, or errors, in patch installations. A common approach for flight software maintenance is to access the original development staff. The argument for utilizing the development staff is that the people who developed the software will be the best people to modify the software on-orbit. However, it can quickly becomes a challenge to obtain the services of these key people. They may no longer be available to the organization. They may have a more urgent job to perform, quite likely on another project under different project management. If they havn't worked on the software for a long time, they may need precious time for refamiliarization to the software, testbeds and tools. Further, a lack of insight into issues related to flight software in its on-orbit environment, may find the developer unprepared for the challenges. The second approach is to train a member of the flight operations team to maintain the spacecraft software. This can prove to be a costly and inflexible solution. The person assigned to this duty may not have enough work to do during a problem free period and may have too much to do when a problem arises. If the person is a talented software engineer, he/she may not enjoy the limited software opportunities available in this position; and may eventually leave for newer technology computer science opportunities. Training replacement flight software personnel can be a difficult and lengthy process. The third approach is to assemble a center of excellence for on-orbit spacecraft software maintenance. Personnel in this specialty center can be managed to support flight software of multiple missions at once. The variety of challenges among a set of on-orbit missions, can result in a dedicated, talented staff which is fully trained and available to support each mission's needs. Such staff are not software developers but are rather spacecraft software systems engineers. The cost to any one mission is extremely low because the software staff works and charges, minimally on missions with no current operations issues; and their professional insight into on-orbit software troubleshooting and maintenance methods ensures low risk, effective and minimal-cost solutions to on-orbit issues.

Shell, Elaine M.↗

A methodology for validating software reliability

A significant problem associated with fault tolerant computer system design is how to insure that there are no embedded software errors, so that an avionics computer system meets the required reliability level. To accomplish this, it is necessary to associate a 'probability of failure' with the operational flight program. It would be more correct to say that the probability of excitation of existing latent design errors within the program is required. In this sense, latent software errors are like latent hardware faults, and techniques that were previously used to measure the probability of failure of hardware due to fault latency can be used to measure the probability of failure of the software. A methodology was developed and applied to a flight control program that was known to operate in a well defined environment. The results indicated that the technique could be used to provide a final validation of the software to a specified reliability level and to evaluate the role of flight test in software validation.

Swern, Frederic L.↗

Measurement, estimation, and prediction of software reliability

Quantitative indices of software reliability are defined, and application of three important indices is indicated: (1) reliability measurement, (2) reliability estimation, and (3) reliability prediction. State of the art techniques for each of these procedures are presented together with considerations of data acquisition. Failure classifications and other documentation for comprehensive software reliability evaluation are described.

Hecht, H.↗

Space Tug laser gyro IMU

A redundant inertial measuring unit (IMU) incorporating six strapdown laser gyros and six accelerometers, arranged so that sensitive axes are normal to the faces of a dodecahedron, provides enhanced reliability with reduced hardware weight. Software monitoring of sensor outputs senses failure of sensors and the system is designed for triple redundancy, with built-in test equipment. Attention is centered on redundancy and fail-safe features, and on the closed-path ring laser gyro arrangement.

Morrison, R.↗

Research in computer science

Various graduate research activities in the field of computer science are reported. Among the topics discussed are: (1) failure probabilities in multi-version software; (2) Gaussian Elimination on parallel computers; (3) three dimensional Poisson solvers on parallel/vector computers; (4) automated task decomposition for multiple robot arms; (5) multi-color incomplete cholesky conjugate gradient methods on the Cyber 205; and (6) parallel implementation of iterative methods for solving linear equations.

Ortega, J. M.↗

CONFIG - Adapting qualitative modeling and discrete event simulation for design of fault management systems

CONFIG is a modeling and simulation tool prototype for analyzing the normal and faulty qualitative behaviors of engineered systems. Qualitative modeling and discrete-event simulation have been adapted and integrated, to support early development, during system design, of software and procedures for management of failures, especially in diagnostic expert systems. Qualitative component models are defined in terms of normal and faulty modes and processes, which are defined by invocation statements and effect statements with time delays. System models are constructed graphically by using instances of components and relations from object-oriented hierarchical model libraries. Extension and reuse of CONFIG models and analysis capabilities in hybrid rule- and model-based expert fault-management support systems are discussed.

Malin, Jane T.↗

Reliability of Fault Tolerant Control Systems

This paper reports Part I of a two part effort, that is intended to delineate the relationship between reliability and fault tolerant control in a quantitative manner. Reliability analysis of fault-tolerant control systems is performed using Markov models. Reliability properties, peculiar to fault-tolerant control systems are emphasized. As a consequence, coverage of failures through redundancy management can be severely limited. It is shown that in the early life of a syi1ein composed of highly reliable subsystems, the reliability of the overall system is affine with respect to coverage, and inadequate coverage induces dominant single point failures. The utility of some existing software tools for assessing the reliability of fault tolerant control systems is also discussed. Coverage modeling is attempted in Part II in a way that captures its dependence on the control performance and on the diagnostic resolution.

Wu, N. Eva↗

Multiscale Static Analysis of Notched and Unnotched Laminates Using the Generalized Method of Cells

The generalized method of cells (GMC) is demonstrated to be a viable micromechanics tool for predicting the deformation and failure response of laminated composites, with and without notches, subjected to tensile and compressive static loading. Given the axial [0], transverse [90], and shear [+45/-45] response of a carbon/epoxy (IM7/977-3) system, the unnotched and notched behavior of three multidirectional layups (Layup 1: [0,45,90,-45](sub 2S), Layup 2: [0,60,0](sub 3S), and Layup 3: [30,60,90,-30, -60](sub 2S)) are predicted under both tensile and compressive static loading. Matrix nonlinearity is modeled in two ways. The first assumes all nonlinearity is due to anisotropic progressive damage of the matrix only, which is modeled, using the multiaxial mixed-mode continuum damage model (MMCDM) within GMC. The second utilizes matrix plasticity coupled with brittle final failure based on the maximum principle strain criteria to account for matrix nonlinearity and failure within the Finite Element Analysis--Micromechanics Analysis Code (FEAMAC) software multiscale framework. Both MMCDM and plasticity models incorporate brittle strain- and stress-based failure criteria for the fiber. Upon satisfaction of these criteria, the fiber properties are immediately reduced to a nominal value. The constitutive response for each constituent (fiber and matrix) is characterized using a combination of vendor data and the axial, transverse, and shear responses of unnotched laminates. Then, the capability of the multiscale methodology is assessed by performing blind predictions of the mentioned notched and unnotched composite laminates response under tensile and compressive loading. Tabulated data along with the detailed results (i.e., stress-strain curves as well as damage evolution states at various ratios of strain to failure) for all laminates are presented.

GMC↗

Standardizing Microprocessor and GPU Radiation Test Approaches

Microprocessor, Graphics Processing Units (GPUs) and DDRx memory devices have emerged as promising next-generation technologies that enables both high performance processing and acceleration of complex algorithms for the latest challenges in human spaceflight, autonomous vehicles and artificial intelligence (AI). The feature sets of these devices offer exponential increases to throughput, calculation capability and system autonomy when compared to legacy flight systems. NASA's Electronic Part and Packaging (NEPP) Program has conducted an investigation into the radiation susceptibility of leading edge devices and process technologies by establishing standardized test approaches. Unlike most discrete devices, these require state of the art test systems to induce specific hardware activity similar to application software, thus allowing the characterization of failure modes within the system. To best characterize the tested part, NEPP eliminates variables that may impact device performance under radiation. Simplification of remaining system-level variables leads to an improved understanding of complex computational devices and their intended applications. The failure modes and error signatures that are recorded during testing are used to determine radiation sensitivity of the semiconductor process and the microcode architecture of the design. This presentation will discuss the test methodology that NASA Electronic Parts and Packaging (NEPP) is working to establish for its microprocessor, GPU and DDRx memory test programs to provide guidance on these devices and their underlying technology, in regards to their potential usage in future space flight systems.

GPU↗

Interfacing LabVIEW With Instrumentation for Electronic Failure Analysis and Beyond

The Laboratory Virtual Instrumentation Engineering Workstation (LabVIEW) software is designed such that equipment and processes related to control systems can be operationally lined and controlled by the use of a computer. Various processes within the failure analysis laboratories of NASA's Kennedy Space Center (KSC) demonstrate the need for modernization and, in some cases, automation, using LabVIEW. An examination of procedures and practices with the Failure Analaysis Laboratory resulted in the conclusion that some device was necessary to elevate the potential users of LabVIEW to an operational level in minimum time. This paper outlines the process involved in creating a tutorial application to enable personnel to apply LabVIEW to their specific projects. Suggestions for furthering the extent to which LabVIEW is used are provided in the areas of data acquisition and process control.

Buchanan, Randy K.↗

Simulation Of Failures And Repairs

Automated Reliability/Availability/Maintainability (ARAM) computer program one of software tools designed to assess candidate architectures of data-management system of the Space Station. Evaluates reliability, availability, and maintainability characteristcs of conceptual system. Uses data representing redundancy and maintainability characteristics of system, and reliability parameters of components of equipment. Design based upon simulation of failures and possible subsequent repairs of each unit of equipment included in system. Analyzes effects of failures and repairs on system and maintains statistics of behavior of system from which results of simulation obtained. Written for IBM PC XT/AT or compatible computer.

Vallone, Antonio↗

Critical design load case fatigue and ultimate failure simulation for a 10-m H-type vertical-axis wind turbine

While previous studies investigating critical VAWT design load cases have focused on large and relatively flexible Darrieus designs, the bulk of current commercial products seeking certification fall in the relatively small, stiff, and H-type configuration, such as the XFlow Energy Corporation turbine that this study compares against. Understanding the critical design load case impacts for both fatigue and ultimate failure for this size and type of VAWT are imperative for certification. The abil

Brownstein, Ian↗

Progressive retry for software error recovery in distributed systems

In this paper, we describe a method of execution retry for bypassing software errors based on checkpointing, rollback, message reordering and replaying. We demonstrate how rollback techniques, previously developed for transient hardware failure recovery, can also be used to recover from software faults by exploiting message reordering to bypass software errors. Our approach intentionally increases the degree of nondeterminism and the scope of rollback when a previous retry fails. Examples from our experience with telecommunications software systems illustrate the benefits of the scheme.

Wang, Yi-Min↗