Search NASA⌕ Search

SEARCH · Search NASA

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Reliability measurement during software development

Measurement of software reliability was carried out during the development of data base software for a multi-sensor tracking system. Every run made during this project was scored as success or failure, and supporting data were collected on forms for further analysis. The failure ratio (number of failures per calendar interval divided by total number of runs) and failure rate (number of failures divided by CPU time for the interval) were found to be consistent measures, on a month-to-month basis as well as from module to module, and therefore considered valid indicators of reliability in this environment. Trend lines could be established from these measurements that provide good visualization of the progress on the job as a whole as well as on individual modules. Over one-half of the observed failures were due to factors associated with the specific run submission rather than with the code proper.

Hecht, H.↗

JANTX1N2970B zener diode

Report evaluates effects of power and temperature overstress on General Semi-conductor and Siemens devices. Excessive failure rates limited testing. Failure modes are described.

Source record↗

Data Applicability of Heritage and New Hardware For Launch Vehicle Reliability Models

Bayesian reliability requires the development of a prior distribution to represent degree of belief about the value of a parameter (such as a component's failure rate) before system specific data become available from testing or operations. Generic failure data are often provided in reliability databases as point estimates (mean or median). A component's failure rate is considered a random variable where all possible values are represented by a probability distribution. The applicability of the generic data source is a significant source of uncertainty that affects the spread of the distribution. This presentation discusses heuristic guidelines for quantifying uncertainty due to generic data applicability when developing prior distributions mainly from reliability predictions.

Al Hassan, Mohammad↗

Orbital performance of communication satellite microwave power amplifiers (MPAs)

This paper presents background data on the performance of microwave power amplifiers (MPAs) used as transmitters in currently operating commercial communication satellites. Specifically aspects of two competing MPA types are discussed. These are well known TWTA (travelling wave tube amplifier) and the SSPA (solid state power amplifier). Extensive in-orbit data has been collected from over 2000 MPAs in 1991 and 1993. The study in 1991 invovlved 75 S/C (spacecraft) covering 463 S/C years. The 1993 'second-look' study encompassed a slightly different population of 72 S/C with 497 S/C years of operation. A surprising result of both studies was that SSPAs, although quite reliable, did not achieve the reliability of TWTAs were one-third more reliable in the 1993 study. This was at C-band with comparable power amplifiers, e.g. 6-16W of RF output power and similar gains. Data at K(sub u)-band is for TWTAs only since there are no SSPAs in the current S/C inventory. The other complementary result was that the projected failure rates used as S/C payload design guidelines were, on average, somewhat higher for TWTAs than the actual failure rates uncovered by this study. SSPA rates were as projected.

Strauss, R.↗

Rare events and Griffiths phases in topological quantum error correction

The performance of quantum error correcting (QEC) codes is often studied under the assumption of spatiotemporally uniform error rates. On the other hand, experimental implementations almost always produce heterogeneous error rates, in either space or time, as a result of effects such as imperfect fabrication and/or cosmic rays. It is therefore important to understand if and how their presence can affect the performance of QEC in qualitative ways. Here, in this work, we study the effects of nonuniform error rates in the representative examples of the 1D repetition code and the 2D toric code, focusing on when they have extended spatiotemporal correlations; these may arise, for instance, from rare events (such as cosmic rays) that temporarily elevate error rates over the entire code patch. These effects can be described in the corresponding statistical mechanics models for decoding, where long-range correlations in the error rates lead to extended rare regions of weaker coupling. For the 1D repetition code where the rare regions are linear, we find two distinct decodable phases: a conventional ordered phase in which logical failure rates decay exponentially with the code distance, and a rare-region dominated Griffiths phase in which failure rates are parametrically larger and decay as a stretched exponential. In particular, the latter phase is present when the error rates in the rare regions are above the bulk threshold. For the 2D toric code where the rare regions are planar, we find no decodable Griffiths phase: rare events which boost error rates above the bulk threshold lead to an asymptotic loss of threshold and failure to decode. Unpacking the failure mechanism implies that techniques for suppressing extended sequences of repeated rare events (which, without intervention, will be statistically present with high probability) will be crucial for QEC with the toric code.

classical statistical mechanics↗

Data Applicability of Heritage and New Hardware for Launch Vehicle System Reliability Models

Many launch vehicle systems are designed and developed using heritage and new hardware. In most cases, the heritage hardware undergoes modifications to fit new functional system requirements, impacting the failure rates and, ultimately, the reliability data. New hardware, which lacks historical data, is often compared to like systems when estimating failure rates. Some qualification of applicability for the data source to the current system should be made. Accurately characterizing the reliability data applicability and quality under these circumstances is crucial to developing model estimations that support confident decisions on design changes and trade studies. This presentation will demonstrate a data-source classification method that ranks reliability data according to applicability and quality criteria to a new launch vehicle. This method accounts for similarities/dissimilarities in source and applicability, as well as operating environments like vibrations, acoustic regime, and shock. This classification approach will be followed by uncertainty-importance routines to assess the need for additional data to reduce uncertainty.

Al Hassan Mohammad↗

Software reliability: Repetitive run experimentation and modeling

A software experiment conducted with repetitive run sampling is reported. Independently generated input data was used to verify that interfailure times are very nearly exponentially distributed and to obtain good estimates of the failure rates of individual errors and demonstrate how widely they vary. This fact invalidates many of the popular software reliability models now in use. The log failure rate of interfailure time was nearly linear as a function of the number of errors corrected. A new model of software reliability is proposed that incorporates these observations.

Nagel, P. M.↗

Single event induced transients in I/O devices - A characterization

The results of single-event upset (SEU) testing performed to evaluate the parametric transients, i.e., amplitude and duration, in several I/O devices, and the impact of these transients are discussed. The failure rate of these devices is dependent on the susceptibility of interconnected devices to the resulting transient change in the output of the I/O device. This failure rate, which is a function of the susceptibility of the interconnected device as well as the SEU response of the I/O device itself, may be significantly different from an upset rate calculated without taking these factors into account. The impact at the system level is discussed by way of an example.

Newberry, D. M.↗

Hierarchical memories: Simulating quantum LDPC codes with local gates

Constant-rate low-density parity-check (LDPC) codes are promising candidates for constructing efficient fault-tolerant quantum memories. However, if physical gates are subject to geometric-locality constraints, it becomes challenging to realize these codes. In this paper, we construct a new family of [[N,K,D]] codes, referred to as hierarchical codes, that encode a number of logical qubits K=Ω(N/log(N) 2 ). The N th element of this code family is obtained by concatenating a constant-rate quantum LDPC code with a surface code; nearest-neighbor gates in two dimensions are sufficient to implement the corresponding syndrome-extraction circuit and achieve a threshold. Below threshold the logical failure rate vanishes superpolynomially as a function of the distance D(N). We present a bilayer architecture for implementing the syndrome-extraction circuit, and estimate the logical failure rate for this architecture. Under conservative assumptions, we find that the hierarchical code outperforms the basic encoding where all logical qubits are encoded in the surface code.

Pattison, Christopher A. [California Institute of ↗

Sensitivity of a critical tracking task to alcohol impairment

A first order critical tracking task is evaluated for its potential to discriminate between sober and intoxicated performances. Mean differences between predrink and postdrink performances as a function of BAC are analyzed. Quantification of the results shows that intoxicated failure rates of 50% for blood alcohol concentrations (BACs) at or above 0.1%, and 75% for BACs at or above 0.14%, can be attained with no sober failure rates. A high initial rate of learning is observed, perhaps due to the very nature of the task whereby the operator is always pushed to his limit, and the scores approach a stable asymptote after approximately 50 trials. Finally, the implementation of the task as an ignition interlock system in the automobile environment is discussed. It is pointed out that lower critical performance limits are anticipated for the mechanized automotive units because of the introduction of larger hardware and neuromuscular lags. Whether such degradation in performance would reduce the effectiveness of the device or not will be determined in a continuing program involving a broader based sample of the driving population and performance correlations with both BACs and driving proficiency.

Tennant, J. A.↗

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability↗

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability↗

Surf-Deformer: Mitigating Dynamic Defects on Surface Code via Adaptive Deformation

In this paper, we introduce Surf-Deformer, a code deformation framework that seamlessly integrates adaptive defect mitigation functionality into the current surface code workflow. It crafts several basic deformation instructions based on fundamental gauge transformations, which can be combined to explore a larger design space than previous methods. This enables more optimized deformation processes tailored to specific defect situations, restoring the QEC capability of deformed codes more efficiently with minimal qubit resources. Additionally, we design an adaptive code layout that accommodates our defect mitigation strategy while ensuring efficient execution of logical operations. Our evaluation shows that Surf-Deformer outperforms previous methods by significantly reducing the end-to-end failure rate of various quantum programs by 35× to 70×, while requiring only about 50% of the qubit resources compared to the previous method to achieve the same level of failure rate. Ablation studies show that Surf-Deformer surpasses previous defect removal methods in preserving QEC capability and facilitates surface code communication by achieving nearly optimal throughput.

Yin, Keyi↗

Risk and Performance Assessment of Generic Mission Architectures: Showcasing the Artemis Mission

A has initiated a strong push to return face. In this work, we astronaut assess performance and risk for proposed mission architectures using a new Mission Architecture Risk Assessment (MARA) tool. The MARA tool can produce statistics about the availability of components and overall performance of the mission considering potential failures of any of its components. In a Monte Carlo approach, the tool repeats the mission simulation multiple times while a random generator lets modules fail according to their failure rates. The results provide statistically meaningful insights into the overall performance of the chosen architecture. A given mission architecture can be freely replicated in the tool, with the mission timeline and basic characteristics of employed mission modules (habitats, rovers, power generation units, etc.) specified in a configuration file. Crucially, failure rates for each module need to be known or estimated. The tool performs an event-driven simulation of the mission and accounts for random failure events. Failed modules can be repaired, which takes crew time but restores operations. In addition to tracking individual modules, MARA can assess the availability of predefined functions throughout the mission. For instance, the function of resource collection would require a rover to collect the resources, a power generation unit to charge the rover, and a resource processing module. Together, the modules that are required for a given function are called a functional group. Similarly, we can assess how much crew time is available to achieve a mission benefit (e.g. research, building a base, etc) as opposed to spending crew time on repairs. Here we employ the method on the proposed NASA Artemis mission. Artemis aims to return United States astronauts to the lunar surface by 2024. Results provide insights into mission failure probabilities, up- and downtime for individual modules and crew-time resources spent on the repair of failed modules. The tool also allows us to tweak the mission architecture in order to find setups that produce more favorable mission performance. As such, the tool can be an aid in improving the mission architect abling cost-benefit analysis for mission improvement.

Rumpf, Clemens M.↗

Impacts of PV Module Connector Failures on Cost and Performance of Utility Scale Photovoltaic Systems

The reliability, cost and performance of electrical connectors are a concern in all types of electrical systems, and demands on connectors used on photovoltaic (PV) systems include that connectors maintain electrical conductivity and physical strength, endure ultraviolet sunlight and high ambient temperature, and resist moisture and chemical intrusion over a very long (>25 year) performance period. Connector failures increase operation and maintenance (O&M) costs and reduce plant production, but connector failure can also cause safety and liability problems, which are of greater concern. This work results from a three-year collaboration between Sandia National Laboratories (SNL), the Electric Power Research Institute (EPRI), and the National Renewable Energy Laboratory (NREL) and funded by the U.S. Department of Energy (DOE) Solar Energy Technology Office (SETO) under Agreements #39035 and #38531 "Connector Reliability Across the US Solar Sector." a multi-pronged investigation of PV connector health across the US (see https://energy.sandia.gov/pvconnectors/). This report presents derivation of a Techno-Economic Analysis (TEA) that models failure modes and frequencies (how often failure occurs), estimates O&M costs and lost production associated with connector failures, and then calculates the effect that PV module connectors can have on Levelized Cost of Energy (LCOE). The model is informed with initial data from quantitative assessment of failure rates, root causes and mechanisms, in-situ diagnostics and data collection, lab-based forensics, and interviews with PV connector manufacturers and plant operators. SNL conducted site inspections at multiple utility-scale sites in different climates and subjected field samples of new, used, and degraded connectors to visual and electrical characterization. EPRI conducted metallurgical analysis of the pin and sleeve conductors to study failure-induced morphological and compositional changes. There is in general a shortage of statistically valid data, but data from PVROM database maintained by SNL was sufficient to ascertain failure rates and lost production as well as provide qualitative insight in its curated maintenance records. This report details the structure of the mathematical model but the sources of data to inform the model will continue to evolve. Analysis of a 100 MW PV plant is provided as an example of the use of the model, with results indicating that connectors are responsible for Annualized O&M Costs of $\$$71,933/year; Annualized Unit O&M Costs of $\$$0.72/kW/year; that a Reserve Account of $\$$187,220 should be available to fund repairs related to connectors; that connectors add $\$$1,494,004 to the Net Present Value of the O&M Costs (project life); and that O&M related to connectors adds about $\$$0.00088/kWh to the Levelized Cost of Energy. The impact of this model is to provide a tool to make the US solar sector more robust by quantifying and monetizing the reliability risks to utility-scale PV systems posed by poorly installed, mismatched and/or poorly designed and manufactured connectors. The TEA provides a model incorporating failure statistics, O&M cost data, and lost production into a single figure of merit, informing decisions and enabling practitioners to optimize cost and performance trade-offs. Stakeholders include connector manufacturers, system designers and equipment specifiers, standards bodies, installers and O&M providers, investors and insurance underwriters. This report supports continued growth of PV predicated on assurances that properly installed and maintained PV system connectors are safe and reliable. The project team is proposing future work including accelerated testing of connectors and expanding the approach taken here to other PV system components, such as TEA for rapid shut-down devices.

14 SOLAR ENERGY↗

NASA Helps Keep the Light Burning for the Saturn Car Company

The Saturn Electronics & Engineering, Inc. (Saturn) facility in Marks, Miss., that produces lamp assemblies was experiencing itermittent problems with its automotive under the hood lamps. After numerous testing and engineering efforts, technicians could not pin down the root of the problem. So Saturn contacted the NASA Technology Assistance Program (TAP) at Stennis Space Center. The Marks production facility had been experiencing intermittent problems with under the hood lamp assemblies for some time. The failure rate, at 2 percent, was unacceptable. Every effort was made to identify the problem so that corrective action could be put in place. The problem was investigated and researched by Saturn's engineering department. In addition, Saturn brought in several independent testing laboratories. Other measures included examining the switch component suppliers and auditing them for compliance to the design specifications and for surface contaminants. All attempts to identify the factors responsible for the failures were inconclusive. In an effort to get to the root of the problem, and at the recommendation of the Mississippi Department of Economic Development, Saturn contacted the NASA TAP at Stennis. The NASA Materials and Contamination Laboratory, with assistance from the Stennis Prototype Laboratory, conducted a materials evaluation study on the switch components. The laboratory findings showed the failures were caused by a build-up of carbon-based contaminants on the switch components. Saturn Electronics & Engineering, Inc., is a minority-owned provider of contract manufacturing services to a diverse global marketplace. Saturn operates manufacturing facilities globally serving the North American, European, and Asian markets. Saturn's production facility in Marks, Mississippi, produces more than 1,000,000 lamps and switches monthly. "Since the NASA recommendations were implemented, our internal failure rate for intermittency has dropped to less than .02 percent. Most importantly, we restored our high-level of customer satisfaction. Stennis provided an invaluable service to our business," Patrick said. Both NASA and Saturn were pleased with the results form this technical assistance project. The Technology Assistance Program at Stennis makes available to the public NASA technical expertise and access to lab facilities. This project provided both services with a positive outcome.

Source record↗

Using Technical Performance Measures

All programs have requirements. For these requirements to be met, there must be a means of measurement. A Technical Performance Measure (TPM) is defined to produce a measured quantity that can be compared to the requirement. In practice, the TPM is often expressed as a maximum or minimum and a goal. Example TPMs for a rocket program are: vacuum or sea level specific impulse (lsp), weight, reliability (often expressed as a failure rate), schedule, operability (turn-around time), design and development cost, production cost, and operating cost. Program status is evaluated by comparing the TPMs against specified values of the requirements. During the program many design decisions are made and most of them affect some or all of the TPMs. Often, the same design decision changes some TPMs favorably while affecting other TPMs unfavorably. The problem then becomes how to compare the effects of a design decision on different TPMs. How much failure rate is one second of specific impulse worth? How many days of schedule is one pound of weight worth? In other words, how to compare dissimilar quantities in order to trade and manage the TPMs to meet all requirements. One method that has been used successfully and has a mathematical basis is Utility Analysis. Utility Analysis enables quantitative comparison among dissimilar attributes. It uses a mathematical model that maps decision maker preferences over the tradeable range of each attribute. It is capable of modeling both independent and dependent attributes. Utility Analysis is well supported in the literature on Decision Theory. It has been used at Pratt & Whitney Rocketdyne for internal programs and for contracted work such as the J-2X rocket engine program. This paper describes the construction of TPMs and describes Utility Analysis. It then discusses the use of TPMs in design trades and to manage margin during a program using Utility Analysis.

Garrett, Christopher J.↗

Management of Risks Associated with Application of Novel Materials in Novel Operating Environments in Novel Reactor Designs

There is currently no widely agreed, detailed general method for licensing a novel plant incorporating novel materials (or materials being deployed in novel environments); in many such situations, there are no directly applicable engineering code cases for decision-makers (including regulators) to rely on. This paper discusses a framework for solving this problem that is based on the Reliability and Integrity Management (RIM) approach delineated in ASME BPVC Section XI Division 2. NRC Regulatory Guide 1.246, Rev. 0, endorses, with conditions, the subject portion of the ASME Code. The proposed framework is meant to support development of a licensing case by addressing certain technical challenges. The framework discussed here is compatible with the Licensing Modernization Project, but applying it in a specific case will call for advances in the state of practice, if not the state of the art. The RIM approach calls for applicants to (a) allocate reliability targets to plant structures, systems, and components (SSCs), (b) show that they are able to relate the currently observed physical condition of each SSC in the program to its failure probability well enough to determine whether the target reliability allocations are being satisfied, allowing for uncertainty related to the novelty of the materials/designs/operating environments, and (c) be able to demonstrate that the proposed program of surveillances will reliably detect unacceptable degradation of an SSC before SSC failure occurs. These challenges are discussed in the paper, and a potentially applicable modeling approach based on cumulative damage rather than failure rates is briefly illustrated.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗