Search NASA⌕ Search

SEARCH · Search NASA

Results for “reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Going Beyond Reliability to Robustness and Resilience in Space Life Support Systems

The words reliability, robustness, and resilience are often used interchangeably to describe tough and dependable systems but the distinctions between them suggest how to design more serviceable space systems. Reliability is simply the quality of consistently performing well. A system that dependably meets its design requirements in the specified environment is reliable. The designers may not consider themselves responsible for failures under unanticipated conditions. Robustness is the capability of performing without failure under a wide range of conditions, which can go beyond the expected range to include possible off-nominal conditions. Resilience is the ability to recover from or adapt to unanticipated damaging events, such as failures, accidents, external disruptions, and repurposing. Such changes can invalidate the usual operating assumptions and cause system failure. Reliability, robustness, and resilience describe dependable performance under increasingly difficult conditions, first the specified environment, then a wider possible environment, and finally unanticipated damaging conditions. These three qualities are increasingly desirable and increasingly difficult to achieve. Engineering for resilience would design systems that can ignore or repair failures, survive accidents, and recover from unanticipated disruptions. Increasing the resilience of space systems would greatly increase space crew safety. Improving reliability and robustness requires dealing with known problems, but improving resilience requires implementing a general approach to reducing the impact of unknown future events. The need for robustness and resilience has been stated for decades but little has been done. Systems designers often assume that they understand everything they need to know. The potential failures caused by changes, failures, accidents, unknown environments, and unknown unknowns can be ignored. Such overconfidence can lead to neglect of reliability, robustness, and resilience.

Harry W. Jones↗

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY↗

2024 Photovoltaic Inverter Reliability Workshop Summary Report & Proceedings

The National Renewable Energy Laboratory (NREL) organized the 2024 Photovoltaic Inverter Reliability Workshop on April 11-12, 2024, hosted at NREL's South Table Mountain campus in Golden, Colorado. The workshop was organized around seven key topics, including the present state of inverter reliability; solutions for reliability challenges; life cycle cost and ownership issues; testing, standards, performance, and reliability metrics; data reporting, analytics, and sharing; and the future of PV inverter reliability research. Participants included inverter manufacturers, national laboratory researchers, academics, independent testing laboratories, and more. Over the course of the two-day workshop, attendees arrived at several key priorities and conclusions. This report summarizes these conclusions and then collects presentations from the workshop into a record of the workshop's proceedings.

14 SOLAR ENERGY↗

Electric Drive Technologies Consortium (EDTC)/ Cost competitive, high-Performance, highly Reliable (CPR) Power Devices on 4H-SiC (Final Report)

4H-Silicon carbide (4H-SiC) is a wide bandgap semiconductor that offers superior material properties over silicon, including higher critical electric field, thermal conductivity, and electron saturation velocity. These advantages make 4H-SiC highly attractive for high-voltage, high-efficiency power electronics. However, realizing the full potential of SiC requires device technologies that are not only high-performing but also manufacturable and reliable under real-world operating conditions. This report summarizes the outcomes of a five-year R&D effort funded by the U.S. Department of Energy (DOE) under the Electric Drive Technologies Consortium (EDTC), focused on developing cost-competitive, high-performance, and highly reliable (CPR) power devices on 4H-SiC substrates. The program targeted scalable and manufacturable 1.2 kV-class SiC MOSFETs optimized for next-generation electric vehicles, renewable energy systems, and industrial power conversion. The project delivered transformative advancements in SiC power device performance and ruggedness. Particularly, Specific on-resistance (R on,sp ) was reduced by up to 37%, from ~4.0 m$\Omega \cdot$cm 2 in earlier designs to an industry-leading 2.40 m$\Omega \cdot$cm 2 , driven by optimized doping, refined JFET widths, and layout engineering. Breakdown voltages (BV) exceeded 1600 V, marking improvement over legacy baselines, and demonstrating the robustness of newly implemented junction profiles and edge terminations. Short-circuit withstand time (SCWT) saw a remarkable 4$\times$ increase, from ~2 $\mu$s to over 8 $\mu$s, achieved through the successful deployment of deep P-well structures (~1.8–2.0 $\mu$m) via channeling implantation. This innovative process breakthrough enabled precise junction formation without MeV-class implantation tools, reduced leakage under high field stress, and allowed even the shortest-channel devices (down to 0.3 $\mu$m) to achieve both high BV and excellent ruggedness—breaking the traditional trade-off between conduction efficiency and blocking capability. Several novel architectures pushed the performance envelope further. JBSFETs—featuring embedded Schottky portions—eliminated bipolar degradation and drastically reduced third-quadrant leakage, while Ladder MOSFETs introduced a clever orthogonal conduction path that achieved a 15.4% reduction in R on,sp over standard linear designs. Switching performance reached new benchmarks: short-channel devices showed a 31% reduction in total switching energy compared to 0.5 $\mu$m counterparts, while maintaining manageable gate drive requirements. Layout-optimized structures not only improved transconductance but also accelerated switching transitions, pointing to real-world benefits in converter-level efficiency. The devices also passed rigorous reliability validation. Stress-tested across TDDB, HTGB, HTRB, HVP, and burn-in, the devices screened under 30 V/10 hr and 43 V/1 s protocols consistently exhibited tighter lifetime distributions and long-term oxide robustness. These screening techniques proved effective in identifying latent defects and ensuring deployment-grade reliability. Meanwhile, advanced 3D TCAD simulations revealed and resolved electric field hotspots—particularly in HEXFET corners—where fields exceeding 4.8 MV/cm were mitigated through geometry-aware layout corrections. Overall, the results of this project demonstrate a manufacturable and scalable SiC power device platform that addresses key DOE performance targets for efficient, robust, and reliable 1.2kV 4H-SiC Power Devices. The developed technologies represent a meaningful step forward in the commercial readiness of high-voltage SiC solutions and provide a strong foundation for continued advancement in wide bandgap power electronics.

42 ENGINEERING↗

Comparison of IC and MEMS packaging reliability approaches

This paper reviews the current status of IC and MEMS packaging technology with emphasis on reliability, compares the norm for IC packaging reliability evaluation and identifies challenges for development of reliability methodologies for MEMS, and finally, proposes the use of COTS MEMS in order to start generating statistically meaningful reliability data as a vehicle for future standardization of reliability test methodology for MEMS packaging.

microelectromechanical systems (MEMS) chip scale p↗

Ultra reliability at NASA

Ultra reliable systems are critical to NASA particularly as consideration is being given to extended lunar missions and manned missions to Mars. NASA has formulated a program designed to improve the reliability of NASA systems. The long term goal for the NASA ultra reliability is to ultimately improve NASA systems by an order of magnitude. The approach outlined in this presentation involves the steps used in developing a strategic plan to achieve the long term objective of ultra reliability. Consideration is given to: complex systems, hardware (including aircraft, aerospace craft and launch vehicles), software, human interactions, long life missions, infrastructure development, and cross cutting technologies. Several NASA-wide workshops have been held, identifying issues for reliability improvement and providing mitigation strategies for these issues. In addition to representation from all of the NASA centers, experts from government (NASA and non-NASA), universities and industry participated. Highlights of a strategic plan, which is being developed using the results from these workshops, will be presented.

risk↗

High-Reliability Pump Module for Non-Planar Ring Oscillator Laser

We propose and have demonstrated a prototype high-reliability pump module for pumping a Non-Planar Ring Oscillator (NPRO) laser suitable for space missions. The pump module consists of multiple fiber-coupled single-mode laser diodes and a fiber array micro-lens array based fiber combiner. The reported Single-Mode laser diode combiner laser pump module (LPM) provides a higher normalized brightness at the combined beam than multimode laser diode based LPMs. A higher brightness from the pump source is essential for efficient NPRO laser pumping and leads to higher reliability because higher efficiency requires a lower operating power for the laser diodes, which in turn increases the reliability and lifetime of the laser diodes. Single-mode laser diodes with Fiber Bragg Grating (FBG) stabilized wavelength permit the pump module to be operated without a thermal electric cooler (TEC) and this further improves the overall reliability of the pump module. The single-mode laser diode LPM is scalable in terms of the number of pump diodes and is capable of combining hundreds of fiber-coupled laser diodes. In the proof-of-concept demonstration, an e-beam written diffractive micro lens array, a custom fiber array, commercial 808nm single mode laser diodes, and a custom NPRO laser head are used. The reliability of the proposed LPM is discussed.

micro lens array↗

FY12 End of Year Report for NEPP DDR2 Reliability

This document reports the status of the NEPP Double Data Rate 2 (DDR2) Reliability effort for FY2012. The task expanded the focus of evaluating reliability effects targeted for device examination. FY11 work highlighted the need to test many more parts and to examine more operating conditions, in order to provide useful recommendations for NASA users of these devices. In order to develop these approaches, it is necessary to develop test capability that can identify reliability outliers. To do this we must test many devices to ensure outliers are in the sample, and we must develop characterization capability to measure many different parameters. For FY12 we increased capability for reliability characterization and sample size. We increased sample size by moving from loose devices to DIMMs with an approximate reduction of 20 to 50 times in terms of per DUT cost. By increasing sample size we have improved our ability to characterize devices that may be considered reliability outliers. This report provides an update on the effort to improve DDR2 testing capability. Although focused on DDR2, the methods being used can be extended to DDR and DDR3 with relative ease.

temperature stress↗

Diverse Redundant Systems for Reliable Space Life Support

Reliable life support systems are required for deep space missions. The probability of a fatal life support failure should be less than one in a thousand in a multi-year mission. It is far too expensive to develop a single system with such high reliability. Using three redundant units would require only that each have a failure probability of one in ten over the mission. Since the system development cost is inverse to the failure probability, this would cut cost by a factor of one hundred. Using replaceable subsystems instead of full systems would further cut cost. Using full sets of replaceable components improves reliability more than using complete systems as spares, since a set of components could repair many different failures instead of just one. Replaceable components would require more tools, space, and planning than full systems or replaceable subsystems. However, identical system redundancy cannot be relied on in practice. Common cause failures can disable all the identical redundant systems. Typical levels of common cause failures will defeat redundancy greater than two. Diverse redundant systems are required for reliable space life support. Three, four, or five diverse redundant systems could be needed for sufficient reliability. One system with lower level repair could be substituted for two diverse systems to save cost.

life support↗

Body of Knowledge (BOK) for Leadless Quad Flat No-Lead/bottom Termination Components (QFN/BTC) Package Trends and Reliability

Bottom terminated components and quad flat no-lead (BTC/QFN) packages have been extensively used by commercial industry for more than a decade. Cost and performance advantages and the closeness of the packages to the boards make them especially unique for radio frequency (RF) applications. A number of high-reliability parts are now available in this style of package configuration. This report presents a summary of literature surveyed and provides a body of knowledge (BOK) gathered on the status of BTC/QFN and their advanced versions of multi-row QFN (MRQFN) packaging technologies. The report provides a comprehensive review of packaging trends and specifications on design, assembly, and reliability. Emphasis is placed on assembly reliability and associated key design and process parameters because they show lower life than standard leaded package assembly under thermal cycling exposures. Inspection of hidden solder joints for assuring quality is challenging and is similar to ball grid arrays (BGAs). Understanding the key BTC/QFN technology trends, applications, processing parameters, workmanship defects, and reliability behavior is important when judicially selecting and narrowing the follow-on packages for evaluation and testing, as well as for the low risk insertion in high-reliability applications.

MLF↗

Evolving Reliability and Maintainability Allocations for NASA Ground Systems

This paper describes the methodology that was developed to allocate reliability and maintainability requirements for the NASA Ground Systems Development and Operations (GSDO) program's subsystems. As systems progressed through their design life cycle and hardware data became available, it became necessary to reexamine the previously derived allocations. Allocating is an iterative process; as systems moved beyond their conceptual and preliminary design phases this provided an opportunity for the reliability engineering team to reevaluate allocations based on updated designs and maintainability characteristics of the components. Trade-offs in reliability and maintainability were essential to ensuring the integrity of the reliability and maintainability analysis. This paper will discuss the value of modifying reliability and maintainability allocations made for the GSDO subsystems as the program nears the end of its design phase.

allocations↗

Evolving Reliability and Maintainability Allocations for NASA Ground Systems

This paper describes the methodology and value of modifying allocations to reliability and maintainability requirements for the NASA Ground Systems Development and Operations (GSDO) program’s subsystems. As systems progressed through their design life cycle and hardware data became available, it became necessary to reexamine the previously derived allocations. This iterative process provided an opportunity for the reliability engineering team to reevaluate allocations as systems moved beyond their conceptual and preliminary design phases. These new allocations are based on updated designs and maintainability characteristics of the components. It was found that trade-offs in reliability and maintainability were essential to ensuring the integrity of the reliability and maintainability analysis. This paper discusses the results of reliability and maintainability reallocations made for the GSDO subsystems as the program nears the end of its design phase.

allocations↗

Evolving Reliability and Maintainability Allocations for NASA Ground Systems

This paper describes the methodology and value of modifying allocations to reliability and maintainability requirements for the NASA Ground Systems Development and Operations (GSDO) programs subsystems. As systems progressed through their design life cycle and hardware data became available, it became necessary to reexamine the previously derived allocations. This iterative process provided an opportunity for the reliability engineering team to reevaluate allocations as systems moved beyond their conceptual and preliminary design phases. These new allocations are based on updated designs and maintainability characteristics of the components. It was found that trade-offs in reliability and maintainability were essential to ensuring the integrity of the reliability and maintainability analysis. This paper discusses the results of reliability and maintainability reallocations made for the GSDO subsystems as the program nears the end of its design phase.

reliability↗

Heroic Reliability Improvement in Manned Space Systems

System reliability can be significantly improved by a strong continued effort to identify and remove all the causes of actual failures. Newly designed systems often have unexpected high failure rates which can be reduced by successive design improvements until the final operational system has an acceptable failure rate. There are many causes of failures and many ways to remove them. New systems may have poor specifications, design errors, or mistaken operations concepts. Correcting unexpected problems as they occur can produce large early gains in reliability. Improved technology in materials, components, and design approaches can increase reliability. The reliability growth is achieved by repeatedly operating the system until it fails, identifying the failure cause, and fixing the problem. The failure rate reduction that can be obtained depends on the number and the failure rates of the correctable failures. Under the strong assumption that the failure causes can be removed, the decline in overall failure rate can be predicted. If a failure occurs at the rate of lambda per unit time, the expected time before the failure occurs and can be corrected is 1/lambda, the Mean Time Before Failure (MTBF). Finding and fixing a less frequent failure with the rate of lambda/2 per unit time requires twice as long, time of 1/(2 lambda). Cutting the failure rate in half requires doubling the test and redesign time and finding and eliminating the failure causes.Reducing the failure rate significantly requires a heroic reliability improvement effort.

life support↗

Analysis and Optimization of Test Plans for Advanced Exploration Systems Reliability and Supportability

Future crewed exploration missions beyond Low Earth Orbit (LEO) will operate farther from Earth and be logistically isolated for longer than any previous human spaceflight mission. Under these conditions, supportability and reliability willbestronger drivers of mission mass and risk than they have been in the past. Items with high failure rates, or uncertain failure rates, can result in high spares mass requirements and/or high risk on deep space missions. Testing is a critical element of system development which provides the opportunity to identify and resolve design issues, defects, or other failure modes before they cause problems during a mission. Reliability growth programs can reduce failure rates by identifying and remove failure modes via design changes, and long-duration life testing can provide valuable data to reduce failure rate estimate uncertainty and verify (to some level of confidence) that components are as reliable as expected. Testing activities take time and resources, however, and must be incorporated into program plans in order to be fully effective. This paper presents an integrated reliability test plan analysis and optimization methodology, which has been used to inform Advanced Exploration Systems (AES) Life Support Systems (LSS) ground test planning for future missions. The methodology determines the optimal number of test units to purchase and allocation of test time –split between reliability growth and uncertainty reduction testing –across a given set of items in order to minimize spares mass for a given mission under constraints on total test cost and schedule. Model outputs also include expected spares mass after testing and the expected number of modifications or refurbishments during testing, both of which can inform program planning. Discussion of the model, conclusions, and future work are also presented.

Testing↗

Analysis and Optimization of Test Plans for Advanced Exploration Systems Reliability and Supportability

Future crewed exploration missions beyond Low Earth Orbit (LEO) will operate farther from Earth and be logistically isolated for longer than any previous human spaceflight mission. Under these conditions, supportability and reliability willbestronger drivers of mission mass and risk than they have been in the past. Items with high failure rates, or uncertain failure rates, can result in high spares mass requirements and/or high risk on deep space missions. Testing is a critical element of system development which provides the opportunity to identify and resolve design issues, defects, or other failure modes before they cause problems during a mission. Reliability growth programs can reduce failure rates by identifying and remove failure modes via design changes, and long-duration life testing can provide valuable data to reduce failure rate estimate uncertainty and verify (to some level of confidence) that components are as reliable as expected. Testing activities take time and resources, however, and must be incorporated into program plans in order to be fully effective. This paper presents an integrated reliability test plan analysis and optimization methodology, which has been used to inform Advanced Exploration Systems (AES) Life Support Systems (LSS) ground test planning for future missions. The methodology determines the optimal number of test units to purchase and allocation of test time –split between reliability growth and uncertainty reduction testing –across a given set of items in order to minimize spares mass for a given mission under constraints on total test cost and schedule. Model outputs also include expected spares mass after testing and the expected number of modifications or refurbishments during testing, both of which can inform program planning. Discussion of the model, conclusions, and future work are also presented.

Testing↗

NASA Physics of Failure (PoF) for Reliability

An item’s reliability or longevity is dependent not only on its design but also on how it is used, manufactured, tested, and the stresses it has or will experience. Stresses include operational and environmental exposures to thermal, voltage, current, age/exposure, mechanical, and radiation mechanisms. Therefore, in reliability analysis, it is important to consider the contributions of all these factors when predicting the failure rates of components. Historically, there has been a reliance on handbook data (e.g., MIL-HDBK-217), but experience has shown that these values and distributions are not representative of actual performance. Therefore, to make more credible reliability and risk assessments for its missions, NASA must transition to estimating likelihoods of failure based on an item’s reliability or longevity factors (or the physical susceptibilities and strengths impacting the design’s performance) has or will experience, whenever possible. To facilitate this transition, a Handbook on Methodology for Physics of Failure Based Reliability Assessments has been developed by NASA to assist in applying physics experiences or experimental physics for empirical analysis and conceptualized physics exposures or theoretical physics for deterministic analysis, to develop and aggregate realistic likelihoods of failure leading to more credible forecasts of item performance and longevity. In addition, since it is NASA’s intention that this document continues to evolve based on community lessons learned and the introduction of new assessment methodologies, NASA is encouraging and appreciates the contributions of current and future authors to maintain and enhance this handbook and its supporting case studies.

PoF↗

Computational materials reliability assessment of hydrogen fueled gas turbine power generation engines

The use of blended fuel sources in land based gas turbine engines drives variations in the resulting operational profile (temperatures and pressures) which can impact engine reliability. Furthermore, variability in the manufacture of components affects the resulting microstructure which directly impacts material performance and reliability. Currently, data-driven models are typically used for maintaining and inspecting fleets of engines. Without explicitly capturing material and operational sources of variability conservatism must be used in developing component-level reliability models. Therefore, there exists an opportunity to use information from materials-scale physics models to better inform reliability modeling and reduce conservatism; the impact is more cost-efficient operation and maintenance of current and future fleets. Specifically, this work establishes a computational framework for evaluating the probabilistic high temperature creep performance of hot-section Ni-based superalloys where uncertainty comes from both microstructural and operational variability. A novel high-fidelity physics model which phenomenologically captures grain-boundary sensitive phenomena has been established. A probabilistic calibration procedure was used to calibrate the model and capture uncertainty in the parameterized model coefficients. A design of experiments methodology was established for identifying informative microstructural digital representations for suitable for forward model evaluation. Results show that training a machine-learning surrogate using this design criteria outperforms random selection of microstructural representations. Finally, two surrogate models were developed: (1) a deterministic surrogate model which predicts the local field response given microstructure, constitutive model parameters, and operating conditions (stress, temperature) and (2) a probabilistic model, where uncertainty comes from constitutive law uncertainty, built using denoising diffusion probabilistic models which samples responses given (1) microstructure and (2) operating conditions. These surrogate models enable partner Siemens Energy to rapidly perform UQ analysis specific to creep deformation across a range of microstructures and operating conditions. The impact is that these ML and physics codes can be used to establish more advanced reliability models for the inspection, servicing, and maintenance of land based gas turbine engines.

36 MATERIALS SCIENCE↗