Search NASASearch

SEARCH · Search NASA

Results for “System reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Heroic Reliability Improvement in Manned Space Systems

System reliability can be significantly improved by a strong continued effort to identify and remove all the causes of actual failures. Newly designed systems often have unexpected high failure rates which can be reduced by successive design improvements until the final operational system has an acceptable failure rate. There are many causes of failures and many ways to remove them. New systems may have poor specifications, design errors, or mistaken operations concepts. Correcting unexpected problems as they occur can produce large early gains in reliability. Improved technology in materials, components, and design approaches can increase reliability. The reliability growth is achieved by repeatedly operating the system until it fails, identifying the failure cause, and fixing the problem. The failure rate reduction that can be obtained depends on the number and the failure rates of the correctable failures. Under the strong assumption that the failure causes can be removed, the decline in overall failure rate can be predicted. If a failure occurs at the rate of lambda per unit time, the expected time before the failure occurs and can be corrected is 1/lambda, the Mean Time Before Failure (MTBF). Finding and fixing a less frequent failure with the rate of lambda/2 per unit time requires twice as long, time of 1/(2 lambda). Cutting the failure rate in half requires doubling the test and redesign time and finding and eliminating the failure causes.Reducing the failure rate significantly requires a heroic reliability improvement effort.

life support

Estimates Of The Orbiter RSI Thermal Protection System Thermal Reliability

In support of the Space Shuttle Orbiter post-flight inspection, structure temperatures are recorded at selected positions on the windward, leeward, starboard and port surfaces. Statistical analysis of this flight data and a non-dimensional load interference (NDLI) method are used to estimate the thermal reliability at positions were reusable surface insulation (RSI) is installed. In this analysis, structure temperatures that exceed the design limit define the critical failure mode. At thirty-three positions the RSI thermal reliability is greater than 0.999999 for the missions studied. This is not the overall system level reliability of the thermal protection system installed on an Orbiter. The results from two Orbiters, OV-102 and OV-105, are in good agreement. The original RSI designs on the OV-102 Orbital Maneuvering System pods, which had low reliability, were significantly improved on OV-105. The NDLI method was also used to estimate thermal reliability from an assessment of TPS uncertainties that was completed shortly before the first Orbiter flight. Results fiom the flight data analysis and the pre-flight assessment agree at several positions near each other. The NDLI method is also effective for optimizing RSI designs to provide uniform thermal reliability on the acreage surface of reusable launch vehicles.

Kolodziej, P.

Sensor Selection and Data Validation for Reliable Integrated System Health Management

For new access to space systems with challenging mission requirements, effective implementation of integrated system health management (ISHM) must be available early in the program to support the design of systems that are safe, reliable, highly autonomous. Early ISHM availability is also needed to promote design for affordable operations; increased knowledge of functional health provided by ISHM supports construction of more efficient operations infrastructure. Lack of early ISHM inclusion in the system design process could result in retrofitting health management systems to augment and expand operational and safety requirements; thereby increasing program cost and risk due to increased instrumentation and computational complexity. Having the right sensors generating the required data to perform condition assessment, such as fault detection and isolation, with a high degree of confidence is critical to reliable operation of ISHM. Also, the data being generated by the sensors needs to be qualified to ensure that the assessments made by the ISHM is not based on faulty data. NASA Glenn Research Center has been developing technologies for sensor selection and data validation as part of the FDDR (Fault Detection, Diagnosis, and Response) element of the Upper Stage project of the Ares 1 launch vehicle development. This presentation will provide an overview of the GRC approach to sensor selection and data quality validation and will present recent results from applications that are representative of the complexity of propulsion systems for access to space vehicles. A brief overview of the sensor selection and data quality validation approaches is provided below. The NASA GRC developed Systematic Sensor Selection Strategy (S4) is a model-based procedure for systematically and quantitatively selecting an optimal sensor suite to provide overall health assessment of a host system. S4 can be logically partitioned into three major subdivisions: the knowledge base, the down-select iteration, and the final selection analysis. The knowledge base required for productive use of S4 consists of system design information and heritage experience together with a focus on components with health implications. The sensor suite down-selection is an iterative process for identifying a group of sensors that provide good fault detection and isolation for targeted fault scenarios. In the final selection analysis, a statistical evaluation algorithm provides the final robustness test for each down-selected sensor suite. NASA GRC has developed an approach to sensor data qualification that applies empirical relationships, threshold detection techniques, and Bayesian belief theory to a network of sensors related by physics (i.e., analytical redundancy) in order to identify the failure of a given sensor within the network. This data quality validation approach extends the state-of-the-art, from red-lines and reasonableness checks that flag a sensor after it fails, to include analytical redundancy-based methods that can identify a sensor in the process of failing. The focus of this effort is on understanding the proper application of analytical redundancy-based data qualification methods for onboard use in monitoring Upper Stage sensors.

Garg, Sanjay

[MaRS Project]

The Space Exploration Division of the Safety and Mission Assurances Directorate is responsible for reducing the risk to Human Space Flight Programs by providing system safety, reliability, and risk analysis. The Risk & Reliability Analysis branch plays a part in this by utilizing Probabilistic Risk Assessment (PRA) and Reliability and Maintainability (R&M) tools to identify possible types of failure and effective solutions. A continuous effort of this branch is MaRS, or Mass and Reliability System, a tool that was the focus of this internship. Future long duration space missions will have to find a balance between the mass and reliability of their spare parts. They will be unable take spares of everything and will have to determine what is most likely to require maintenance and spares. Currently there is no database that combines mass and reliability data of low level space-grade components. MaRS aims to be the first database to do this. The data in MaRS will be based on the hardware flown on the International Space Stations (ISS). The components on the ISS have a long history and are well documented, making them the perfect source. Currently, MaRS is a functioning excel workbook database; the backend is complete and only requires optimization. MaRS has been populated with all the assemblies and their components that are used on the ISS; the failures of these components are updated regularly. This project was a continuation on the efforts of previous intern groups. Once complete, R&M engineers working on future space flight missions will be able to quickly access failure and mass data on assemblies and components, allowing them to make important decisions and tradeoffs.

Aruljothi, Arunvenkatesh

Probabilistic Risk-Based Approach to Aeropropulsion System Assessment Developed

In an era of shrinking development budgets and resources, where there is also an emphasis on reducing the product development cycle, the role of system assessment, performed in the early stages of an engine development program, becomes very critical to the successful development of new aeropropulsion systems. A reliable system assessment not only helps to identify the best propulsion system concept among several candidates, it can also identify which technologies are worth pursuing. This is particularly important for advanced aeropropulsion technology development programs, which require an enormous amount of resources. In the current practice of deterministic, or point-design, approaches, the uncertainties of design variables are either unaccounted for or accounted for by safety factors. This could often result in an assessment with unknown and unquantifiable reliability. Consequently, it would fail to provide additional insight into the risks associated with the new technologies, which are often needed by decisionmakers to determine the feasibility and return-on-investment of a new aircraft engine.

Tong, Michael T.

Hydroclimate-coupled framework for assessing power system resilience under summer drought and climate change

Extreme drought, exacerbated by climate change, increasingly threatens power system resilience, and a systematic assessment of such impacts is challenging due to the unpredictability of drought and their associated modeling complexity. Here, to address the challenge, this research develops a hydroclimate-coupled power system resilience assessment framework that enables systematic modeling of drought and climate change impacts on generation, transmission, and demand sectors. Applying the framework to the 2025 Eastern U.S. power grid — comprising 6,055 at-risk generators — under climate-induced summer drought scenarios (including SSP126, SSP245, SSP370, and SSP585) from 2023 to 2100, the study finds that climate-induced droughts could jeopardize the power system’s reliability to a greater extent than historical events, potentially leading to widespread load shedding. More specifically, the study reveals that under the twenty-one representative drought scenarios, the loss of load expectation (LOLE) of the grid could range from 34.77 to 91.48 days per summer. The simulations indicate that implementing resilience enhancement strategies is crucial to ensure reliable system operation, which encompasses initiatives such as demand response, upgrading open cooling systems, and transmission expansion. In all, these findings underscore the urgent need for proactive planning and investment in resilient U.S. power systems to mitigate the impacts of extreme drought events.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Predictions for Radiation Shielding Materials

Radiation from galactic cosmic rays (GCR) and solar particle events (SPE) is a serious hazard to humans and electronic instruments during space travel, particularly on prolonged missions outside the Earth s magnetic fields. Galactic cosmic radiation (GCR) is composed of approx. 98% nucleons and approx. 2% electrons and positrons. Although cosmic ray heavy ions are 1-2% of the fluence, these energetic heavy nuclei (HZE) contribute 50% of the long-term dose. These unusually high specific ionizations pose a significant health hazard acting as carcinogens and also causing microelectronics damage inside spacecraft and high-flying aircraft. These HZE ions are of concern for radiation protection and radiation shielding technology, because gross rearrangements and mutations and deletions in DNA are expected. Calculations have shown that HZE particles have a strong preference for interaction with light nuclei. The best shield for this radiation would be liquid hydrogen, which is totally impractical. For this reason, hydrogen-containing polymers make the most effective practical shields. Shielding is required during missions in Earth orbit and possibly for frequent flying at high altitude because of the broad GCR spectrum and during a passage into deep space and LunarMars habitation because of the protracted exposure encountered on a long space mission. An additional hazard comes from solar particle events (SPEs) which are mostly energetic protons that can produce heavy ion secondaries as well as neutrons in materials. These events occur at unpredictable times and can deliver a potentially lethal dose within several hours to an unshielded human. Radiation protection for humans requires safety in short-term missions and maintaining career exposure limits within acceptable levels on future long-term exploration missions. The selection of shield materials can alter the protection of humans by an order of magnitude. If improperly selected, shielding materials can actually increase radiation damage due to penetration properties and nuclear fragmentation. Protecting space-borne microelectronics from single event upsets (SEUs) by transmitted radiation will benefit system reliability and system design cost by using optimal shield materials. Long-term missions on the surface of the Moon or Mars will require the construction of habitats to protect humans during their stay. One approach to the construction is to make structural materials from lunar or Martian regolith using a polymeric material as a binder. The hydrogen-containing polymers are considerably more effective for radiation protection than the regolith, but the combination minimizes the amount of polymer to be transported. We have made composites of simulated lunar regolith with two different polymers, LaRC-SI, a high-performance polyimide thermoset, and polyethylene, a thermoplastic.

Kiefer, Richard L.

A NASA initiative: Software engineering for reliable complex systems

The objective is the development of methods, technology, and skills that will enable NASA to cost-effectively specify, build, and manage reliable software which can evolve and be maintained over an extended period. The need for such software is rooted in the increasing integration of software and computing components into NASA systems. Current NASA Software Engineering expertise was applied toward some of the largest reliable systems including: shuttle launch; ground support; shuttle simulation; minor control; satellite tracking; and scientific data systems. Unfortunately, no theory exists for reliable complex software systems. NASA is seeking to fill this theoretical gap through a number of approaches. One such approach is to conduct research on theoretical foundations for managing complex software systems. It includes: communication models, new and modified paradigms, and life-cycle models. Another approach is research in the theoretical foundations for reliable software development and validation. It focuses upon formal specifications, programming languages, software engineering systems, software reuse, formal verification, and software safety. Further approaches involve benchmarking a NASA software environment, experimentation within the NASA context, evolution of present NASA methodology, and transfer of technology to the space station software support environment.

Holcomb, Lee B.

An integrated approach to system design, reliability, and diagnosis

The requirement for ultradependability of computer systems in future avionics and space applications necessitates a top-down, integrated systems engineering approach for design, implementation, testing, and operation. The functional analyses of hardware and software systems must be combined by models that are flexible enough to represent their interactions and behavior. The information contained in these models must be accessible throughout all phases of the system life cycle in order to maintain consistency and accuracy in design and operational decisions. One approach being taken by researchers at Ames Research Center is the creation of an object-oriented environment that integrates information about system components required in the reliability evaluation with behavioral information useful for diagnostic algorithms. Procedures have been developed at Ames that perform reliability evaluations during design and failure diagnoses during system operation. These procedures utilize information from a central source, structured as object-oriented fault trees. Fault trees were selected because they are a flexible model widely used in aerospace applications and because they give a concise, structured representation of system behavior. The utility of this integrated environment for aerospace applications in light of our experiences during its development and use is described. The techniques for reliability evaluation and failure diagnosis are discussed, and current extensions of the environment and areas requiring further development are summarized.

Patterson-Hine, F. A.

Techniques for generating highly reliable redundant systems.

A simple heuristic algorithm for designing highly reliable modularly redundant computer systems under complexity constraints is presented. The technique, which produces near optimal solutions, is intuitively appealing and easy to apply. The algorithms performance is shown to compare very well with the optimal solution obtained via a computerized model for dynamic programming.

White, J. B.

HiRel: Hybrid Automated Reliability Predictor (HARP) integrated reliability tool system, (version 7.0). Volume 2: HARP tutorial

The Hybrid Automated Reliability Predictor (HARP) integrated Reliability (HiRel) tool system for reliability/availability prediction offers a toolbox of integrated reliability/availability programs that can be used to customize the user's application in a workstation or nonworkstation environment. The Hybrid Automated Reliability Predictor (HARP) tutorial provides insight into HARP modeling techniques and the interactive textual prompting input language via a step-by-step explanation and demonstration of HARP's fault occurrence/repair model and the fault/error handling models. Example applications are worked in their entirety and the HARP tabular output data are presented for each. Simple models are presented at first with each succeeding example demonstrating greater modeling power and complexity. This document is not intended to present the theoretical and mathematical basis for HARP.

Rothmann, Elizabeth

Ultracapacitor-Based Uninterrupted Power Supply System

The ultracapacitor-based uninterrupted power supply (UPS) system enhances system reliability; reduces life-of-system, maintenance, and downtime costs; and greatly reduces environmental impact when compared to conventional UPS energy storage systems. This design provides power when required and absorbs power when required to smooth the system load and also has excellent low-temperature performance. The UPS used during hardware tests at Glenn is an efficient, compact, maintenance-free, rack-mount, pure sine-wave inverter unit. The UPS provides a continuous output power up to 1,700 W with a surge rating of 1,870 W for up to one minute at a nominal output voltage of 115 VAC. The ultracapacitor energy storage system tested in conjunction with the UPS is rated at 5.8 F. This is a bank of ten symmetric ultracapacitor modules. Each module is actively balanced using a linear voltage balancing technique in which the cell-to-cell leakage is dependent upon the imbalance of the individual cells. The ultracapacitors are charged by a DC power supply, which can provide up to 300 VDC at 4 A. A constant-voltage, constant-current power supply was selected for this application. The long life of ultracapacitors greatly enhances system reliability, which is significant in critical applications such as medical power systems and space power systems. The energy storage system can usually last longer than the application, given its 20-year life span. This means that the ultracapacitors will probably never need to be replaced and disposed of, whereas batteries require frequent replacement and disposal. The charge-discharge efficiency of rechargeable batteries is approximately 50 percent, and after some hundreds of charges and discharges, they must be replaced. The charge-discharge efficiency of ultracapacitors exceeds 90 percent, and can accept more than a million charges and discharges. Thus, there is a significant energy savings through the efficiency improvement, and there is far less downtime for applications and labor involved in replacing an ultracapacitor versus batteries. Also, the lengthy lifespan of this design would greatly reduce the disposal problems posed by lead acid, nickel cadmium, lithium, and nickel metal hydride batteries. This innovation is recyclable by nature, which further reduces system costs. The disposal of ultracapacitors is simple, as they are constructed of non-hazardous components. They are also safer than batteries in that they can be easily discharged, and left indefinitely in a safe, discharged state where batteries cannot.

Eichenberg, Dennis J.

HiRel: Hybrid Automated Reliability Predictor (HARP) integrated reliability tool system, (version 7.0). Volume 1: HARP introduction and user's guide

The Hybrid Automated Reliability Predictor (HARP) integrated Reliability (HiRel) tool system for reliability/availability prediction offers a toolbox of integrated reliability/availability programs that can be used to customize the user's application in a workstation or nonworkstation environment. HiRel consists of interactive graphical input/output programs and four reliability/availability modeling engines that provide analytical and simulative solutions to a wide host of reliable fault-tolerant system architectures and is also applicable to electronic systems in general. The tool system was designed to be compatible with most computing platforms and operating systems, and some programs have been beta tested, within the aerospace community for over 8 years. Volume 1 provides an introduction to the HARP program. Comprehensive information on HARP mathematical models can be found in the references.

Bavuso, Salvatore J.

The Use of Efficient Broadcast Protocols in Asynchronous Distributed Systems

Reliable broadcast protocols are important tools in distributed and fault-tolerant programming. They are useful for sharing information and for maintaining replicated data in a distributed system. However, a wide range of such protocols has been proposed. These protocols differ in their fault tolerance and delivery ordering characteristics. There is a tradeoff between the cost of a broadcast protocol and how much ordering it provides. It is, therefore, desirable to employ protocols that support only a low degree of ordering whenever possible. This dissertation presents techniques for deciding how strongly ordered a protocol is necessary to solve a given application problem. It is shown that there are two distinct classes of application problems: problems that can be solved with efficient, asynchronous protocols, and problems that require global ordering. The concept of a linearization function that maps partially ordered sets of events to totally ordered histories is introduced. How to construct an asynchronous implementation that solves a given problem if a linearization function for it can be found is shown. It is proved that in general the question of whether a problem has an asynchronous solution is undecidable. Hence there exists no general algorithm that would automatically construct a suitable linearization function for a given problem. Therefore, an important subclass of problems that have certain commutativity properties are considered. Techniques for constructing asynchronous implementations for this class are presented. These techniques are useful for constructing efficient asynchronous implementations for a broad range of practical problems.

Schmuck, Frank Bernhard

On reliable control system designs

A mathematical model for use in the design of reliable multivariable control systems is discussed with special emphasis on actuator failures and necessary actuator redundancy levels. The model consists of a linear time invariant discrete time dynamical system. Configuration changes in the system dynamics are governed by a Markov chain that includes transition probabilities from one configuration state to another. The performance index is a standard quadratic cost functional, over an infinite time interval. The actual system configuration can be deduced with a one step delay. The calculation of the optimal control law requires the solution of a set of highly coupled Riccati-like matrix difference equations. Results can be used for off-line studies relating the open loop dynamics, required performance, actuator mean time to failure, and functional or identical actuator redundancy, with and without feedback gain reconfiguration strategies.

Birdwell, J. D.