Research on failure free systems Final report
Fortran IV programming of test point allocations and reliability analysis procedure for redundant digital system, and application to Mariner C sequencer
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Fortran IV programming of test point allocations and reliability analysis procedure for redundant digital system, and application to Mariner C sequencer
Spacecraft human life support systems can achieve ultra reliability by providing sufficient spares to replace all failed components. The additional mass of spares for ultra reliability is approximately equal to the original system mass, provided that the original system reliability is not too low. Acceptable reliability can be achieved for the Space Shuttle and Space Station by preventive maintenance and by replacing failed units. However, on-demand maintenance and repair requires a logistics supply chain in place to provide the needed spares. In contrast, a Mars or other long space mission must take along all the needed spares, since resupply is not possible. Long missions must achieve ultra reliability, a very low failure rate per hour, since they will take years rather than weeks and cannot be cut short if a failure occurs. Also, distant missions have a much higher mass launch cost per kilogram than near-Earth missions. Achieving ultra reliable spacecraft life support systems with acceptable mass will require a well-planned and extensive development effort. Analysis must determine the reliability requirement and allocate it to subsystems and components. Ultra reliability requires reducing the intrinsic failure causes, providing spares to replace failed components and having "graceful" failure modes. Technologies, components, and materials must be selected and designed for high reliability. Long duration testing is needed to confirm very low failure rates. Systems design should segregate the failure causes in the smallest, most easily replaceable parts. The system must be designed, developed, integrated, and tested with system reliability in mind. Maintenance and reparability of failed units must not add to the probability of failure. The overall system must be tested sufficiently to identify any design errors. A program to develop ultra reliable space life support systems with acceptable mass should start soon since it must be a long term effort.
This analysis considers the optimum allocation of redundancy in a system of serially connected subsystems in which each subsystem is of the k-out-of-n type. Redundancy is optimally allocated when: (1) reliability is maximized for given costs; or (2) costs are minimized for given reliability. Several techniques are presented for achieving optimum allocation and their relative merits are discussed. Approximate solutions in closed form were attainable only for the special case of series-parallel systems and the efficacy of these approximations is discussed.
The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.
The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.
Recycling life support systems can achieve ultra reliability by using spares to replace failed components. The added mass for spares is approximately equal to the original system mass, provided the original system reliability is not very low. Acceptable reliability can be achieved for the space shuttle and space station by preventive maintenance and by replacing failed units, However, this maintenance and repair depends on a logistics supply chain that provides the needed spares. The Mars mission must take all the needed spares at launch. The Mars mission also must achieve ultra reliability, a very low failure rate per hour, since it requires years rather than weeks and cannot be cut short if a failure occurs. Also, the Mars mission has a much higher mass launch cost per kilogram than shuttle or station. Achieving ultra reliable space life support with acceptable mass will require a well-planned and extensive development effort. Analysis must define the reliability requirement and allocate it to subsystems and components. Technologies, components, and materials must be designed and selected for high reliability. Extensive testing is needed to ascertain very low failure rates. Systems design should segregate the failure causes in the smallest, most easily replaceable parts. The systems must be designed, produced, integrated, and tested without impairing system reliability. Maintenance and failed unit replacement should not introduce any additional probability of failure. The overall system must be tested sufficiently to identify any design errors. A program to develop ultra reliable space life support systems with acceptable mass must start soon if it is to produce timely results for the moon and Mars.
Mars Sample Return is our Grand Challenge for the coming decade. TPS (Thermal Protection System) nominal performance is not the key challenge. The main difficulty for designers is the need to verify unprecedented reliability for the entry system: current guidelines for prevention of backward contamination require that the probability of spores larger than 1 micron diameter escaping into the Earth environment be lower than 1 million for the entire system, and the allocation to TPS would be more stringent than that. For reference, the reliability allocation for Orion TPS is closer to 11000, and the demonstrated reliability for previous human Earth return systems was closer to 1100. Improving reliability by more than 3 orders of magnitude is a grand challenge indeed. The TPS community must embrace the possibility of new architectures that are focused on reliability above thermal performance and mass efficiency. MSR (Mars Sample Return) EEV (Earth Entry Vehicle) will be hit with MMOD (Micrometeoroid and Orbital Debris) prior to reentry. A chute-less aero-shell design which allows for self-righting shape was baselined in prior MSR studies, with the assumption that a passive system will maximize EEV robustness. Hence the aero-shell along with the TPS has to take ground impact and not break apart. System verification will require testing to establish ablative performance and thermal failure but also testing of damage from MMOD, and structural performance at ground impact. Mission requirements will demand analysis, testing and verification that are focused on establishing reliability of the design. In this proposed talk, we will focus on the grand challenge of MSR EEV TPS and the need for innovative approaches to address challenges in modeling, testing, manufacturing and verification.
Mariner Mars 1964 failure rates computed for use in reliability predictions and cost allocations
This report is a part of the deliverables for technical assistance provided to Cat Creek Energy for the Coral Summit Trybrid (Triple Hybrid-Pumped Storage Hydropower, Battery Energy Storage System, and photovoltaic solar energy) energy project. This report explores the operational benefits and challenges of hybridizing an open loop PSH (200MW) located at Mackay, Custer County, Idaho with solar PV (Ground mount 300MW and floating 40MW) and battery (720MWhr). This document reports two activities performed as a part of the hybridization assessment task 1) optimal resource allocation and energy management strategy, and 2) power quality and reliability assessment. From optimal resource allocation and energy management strategy (activity 1), the following key findings can be observed: • Conventional PSH (CPSH) with two reversible pump turbines and separate penstocks can provide required flexibility equivalent to that from two ternary PSH with separate penstock. With single unit CPSH, upper reservoir head cannot be maintained accurately, the variation of water level is rapid and pump mode flexibility is not available. These disadvantages can be overcome by single unit TPSH. However, using two CPSH units with separate penstocks also overcome these disadvantages with the formation of the hydraulic short circuit between two conventional units. • Flooding of the lower reservoir is a severe concern when considering continuous operation for black start. This limits the duration of continuous operation from PSH alone to around 50 hours. Due to the complementary PV and battery action, the duration of continuous operation and smooth power output can be extended. • An optimization problem is framed that maximizes the power output on an hourly basis while minimizing constraint violations and respecting seasonal variations of solar PV and load profiles . Two value streams, arbitrage and baseload generation are served by this profile. It was uncovered that for smooth power output during regular operation, PV curtailment will be required, or the battery capacity needs to be increased above 90MW to accommodate additional PV. From power quality and reliability assessment (activity 2) the following takeaway points can be observed: • The Trybrid, when integrated at the Lost River bus, and limited to 250MW in generation mode and -150MW in the pump mode, causes no violation of voltage or flow.
High reliability is desired in all engineered systems. One way to improve system reliability is to use redundant components. When redundant components are used, the problem becomes one of allocating them to achieve the best reliability without exceeding other design constraints such as cost, weight, or volume. Systems with few components can be optimized by simply examining every possible combination but the number of combinations for most systems is prohibitive. A computerized iteration of the process is possible but anything short of a super computer requires too much time to be practical. Many researchers have derived mathematical formulations for calculating the optimum configuration directly. However, most of the derivations are based on continuous functions whereas the real system is composed of discrete entities. Therefore, these techniques are approximations of the true optimum solution. This paper describes a computer program that will determine the optimum configuration of a system of multiple redundancy of both standard and optional components. The algorithm is a pair-wise comparative progression technique which can derive the true optimum by calculating only a small fraction of the total number of combinations. A designer can quickly analyze a system with this program on a personal computer.
Results of computer simulation studies of the hybrid pull-up bootstrap decoding algorithm, using a constraint length 24, nonsystematic, rate 1/2 convolutional code for the symmetric channel with both binary and eight-level quantized outputs. Computational performance was used to measure the effect of several decoder parameters and determine practical operating constraints. Results reveal that the track length may be reduced to 500 information bits with small degradation in performance. The optimum number of tracks per block was found to be in the range from 7 to 11. An effective technique was devised to efficiently allocate computational effort and identify reliably decoded data sections. Long simulations indicate that a practical bootstrap decoding configuration has a computational performance about 1.0 dB better than sequential decoding and an output bit error rate about .0000025 near the R sub comp point.
Future crewed exploration missions beyond Low Earth Orbit (LEO) will operate farther from Earth and be logistically isolated for longer than any previous human spaceflight mission. Under these conditions, supportability and reliability willbestronger drivers of mission mass and risk than they have been in the past. Items with high failure rates, or uncertain failure rates, can result in high spares mass requirements and/or high risk on deep space missions. Testing is a critical element of system development which provides the opportunity to identify and resolve design issues, defects, or other failure modes before they cause problems during a mission. Reliability growth programs can reduce failure rates by identifying and remove failure modes via design changes, and long-duration life testing can provide valuable data to reduce failure rate estimate uncertainty and verify (to some level of confidence) that components are as reliable as expected. Testing activities take time and resources, however, and must be incorporated into program plans in order to be fully effective. This paper presents an integrated reliability test plan analysis and optimization methodology, which has been used to inform Advanced Exploration Systems (AES) Life Support Systems (LSS) ground test planning for future missions. The methodology determines the optimal number of test units to purchase and allocation of test time –split between reliability growth and uncertainty reduction testing –across a given set of items in order to minimize spares mass for a given mission under constraints on total test cost and schedule. Model outputs also include expected spares mass after testing and the expected number of modifications or refurbishments during testing, both of which can inform program planning. Discussion of the model, conclusions, and future work are also presented.
Future crewed exploration missions beyond Low Earth Orbit (LEO) will operate farther from Earth and be logistically isolated for longer than any previous human spaceflight mission. Under these conditions, supportability and reliability willbestronger drivers of mission mass and risk than they have been in the past. Items with high failure rates, or uncertain failure rates, can result in high spares mass requirements and/or high risk on deep space missions. Testing is a critical element of system development which provides the opportunity to identify and resolve design issues, defects, or other failure modes before they cause problems during a mission. Reliability growth programs can reduce failure rates by identifying and remove failure modes via design changes, and long-duration life testing can provide valuable data to reduce failure rate estimate uncertainty and verify (to some level of confidence) that components are as reliable as expected. Testing activities take time and resources, however, and must be incorporated into program plans in order to be fully effective. This paper presents an integrated reliability test plan analysis and optimization methodology, which has been used to inform Advanced Exploration Systems (AES) Life Support Systems (LSS) ground test planning for future missions. The methodology determines the optimal number of test units to purchase and allocation of test time –split between reliability growth and uncertainty reduction testing –across a given set of items in order to minimize spares mass for a given mission under constraints on total test cost and schedule. Model outputs also include expected spares mass after testing and the expected number of modifications or refurbishments during testing, both of which can inform program planning. Discussion of the model, conclusions, and future work are also presented.
The control allocation problem was investigated for linear dynamical systems with known parametric uncertainty. Minimizing a cost function that penalizes the variance of the error in achieving commanded forces and moments on the vehicle resulted in a special case of the weighted pseudo-inverse allocator. Rather than an engineer designing the weighting matrix, it is computed from the covariances of the control effectiveness parameters. This minimum-variance allocator balances the effectiveness of the control inputs against the corresponding levels of uncertainty. The approach was demonstrated using simulations of aircraft with realistic uncertainty levels operating in open-loop and closed-loop configurations. Results showed that when model uncertainty is known, significant, and unevenly distributed amongst the controls, the minimum-variance allocator more often achieves the intended forces and moments on the vehicle in comparison to other allocators, which can lead to increased performance, reliability, and safety during flight tests. The cost for this robustness is a diminished achievable moment space for the vehicle.
A probabilistic resource allocation system is disclosed containing a low capacity computational module (Short Term Memory or STM) and a self-organizing associative network (Long Term Memory or LTM) where nodes represent elementary resources, terminal end nodes represent goals, and weighted links represent the order of resource association in different allocation episodes. Goals and their priorities are indicated by the user, and allocation decisions are made in the STM, while candidate associations of resources are supplied by the LTM based on the association strength (reliability). Weights are automatically assigned to the network links based on the frequency and relative success of exercising those links in the previous allocation decisions. Accumulation of allocation history in the form of an associative network in the LTM reduces computational demands on subsequent allocations. For this purpose, the network automatically partitions itself into strongly associated high reliability packets, allowing fast approximate computation and display of allocation solutions satisfying the overall reliability and other user-imposed constraints. System performance improves in time due to modification of network parameters and partitioning criteria based on the performance feedback.
A probabilistic resource allocation system is disclosed containing a low capacity computational module (Short Term Memory or STM) and a self-organizing associative network (Long Term Memory or LTM) where nodes represent elementary resources, terminal end nodes represent goals, and directed links represent the order of resource association in different allocation episodes. Goals and their priorities are indicated by the user, and allocation decisions are made in the STM, while candidate associations of resources are supplied by the LTM based on the association strength (reliability). Reliability values are automatically assigned to the network links based on the frequency and relative success of exercising those links in the previous allocation decisions. Accumulation of allocation history in the form of an associative network in the LTM reduces computational demands on subsequent allocations. For this purpose, the network automatically partitions itself into strongly associated high reliability packets, allowing fast approximate computation and display of allocation solutions satisfying the overall reliability and other user-imposed constraints. System performance improves in time due to modification of network parameters and partitioning criteria based on the performance feedback.
Laser data relays potentially offer continuous 1 Gb/sec bandwidths, drastically increasing low-altitude satellite data collection capacity over present store-and-dump techniques. Availability of the laser link as a reliable alternative, operating within conventional low-altitude communication subsystem weight and power allocations, will create customer pressure for adoption. Major communication relay system impacts are discussed including reliability, mechanical design, attitude control, on-board data handling, contamination control, and traffic-net management. Interface parameters which drive the fundamental relay satellite design concepts are discussed, and conditions requiring early quantitative analysis are identified.
Despite their crucial role in supplying heat and power to universities, industries, and healthcare facilities, many steam-based district heating systems rely on outdated control methods. Among these, multi-central plant districts are particularly challenging due to the complexities of coordinating multiple plants, optimizing load distributions, and managing system downtime. In response, new operational strategies are developed to enhance the efficiency and sustainability of steam districts while utilizing existing resources. These strategies include reducing plant operational pressure without compromising the reliable supply to buildings and optimizing load allocation across multiple plants. The load allocation considers boiler part-load efficiency, runtime, network losses, and building pressure set points, and is compared with traditional multi-boiler controls. To support this exploration, new dynamic Modelica models are developed. In addition, methods to reduce modeling complexities are incorporated, enhancing their suitability for practical applications. A holistic district-wide analysis using a real university case study demonstrates a 4.7% fuel savings by lowering boiler operational pressure from 900 kPa to 600 kPa, along with a 13.3% reduction in condensation losses across the distribution network. Furthermore, the load allocation approach results in a 13.1% reduction in fuel consumption during peak winter periods and 15.3% during shoulder periods, with corresponding decreases in carbon emissions and fuel costs. This approach can also save maintenance costs by reducing the boiler runtime by 49.6%. In conclusion, this research underscores the benefits of retrofitting aging steam district heating systems, offering immediate operational improvements by enhancing efficiency, meeting regulatory compliance, and extending infrastructure lifespans while delaying costly overhauls.