Search NASASearch

SEARCH · Search NASA

Results for “System reliability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

General Aviation Aircraft Reliability Study

This reliability study was performed in order to provide the aviation community with an estimate of Complex General Aviation (GA) Aircraft System reliability. To successfully improve the safety and reliability for the next generation of GA aircraft, a study of current GA aircraft attributes was prudent. This was accomplished by benchmarking the reliability of operational Complex GA Aircraft Systems. Specifically, Complex GA Aircraft System reliability was estimated using data obtained from the logbooks of a random sample of the Complex GA Aircraft population.

Pettit, Duane

Aerospace Applications of Weibull and Monte Carlo Simulation with Importance Sampling

Recent developments in reliability modeling and computer technology have made it practical to use the Weibull time to failure distribution to model the system reliability of complex fault-tolerant computer-based systems. These system models are becoming increasingly popular in space systems applications as a result of mounting data that support the decreasing Weibull failure distribution and the expectation of increased system reliability. This presentation introduces the new reliability modeling developments and demonstrates their application to a novel space system application. The application is a proposed guidance, navigation, and control (GN&C) system for use in a long duration manned spacecraft for a possible Mars mission. Comparisons to the constant failure rate model are presented and the ramifications of doing so are discussed.

Bavuso, Salvatore J.

A Probabilistic Approach of Incorporating Safety and Reliability in System Designs for a Manned Mission to Mars

Conceptual stages in mission design often lack the input of quantitative safety and reliability assessments, simply because failure rates or other data are not yet available for systems that have not yet been designed. Absence of such data should not, however, prevent the development of a quantitative risk models with placeholders for missing data. Functions (that is, actions the systems must perform) in mission design will eventually require system probabilities of success, and there could be much learned from surrogate data, adequately bounded in uncertainty, used in a large event tree model of a complex mission.

Railsback, Jan W.

Automating Anomaly Detection for Target systems at Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory, produces the world’s most intense pulse neutrons beams. An accelerated proton beam is directed into a mercury target to generate neutrons via spallation. The target system accounted for over 40% of the overall downtime of the facility in 2022. Thus, early detection in anomalies in the target systems can enable taking corrective actions to avoid failures and reduce downtime. Fault prognostics and anomaly detection in accelerators, both at SNS and outside, has largely focused on the beam side. This paper presents one the first studies exploring leveraging machine learning to automate the detection of anomalies in the target system. The target system consists of over 30 different interconnected subsystems, and the present work focuses on the mercury process system as a use case. Analyzing data from 28 process variables from 2022 and 2023, tree-based and reconstruction-based algorithms are employed to detect anomalies in archived data. The algorithms detected previously unreported anomalies, several of which were deemed alert worthy by human experts, particularly those found by reconstruction-based algorithms. Using data from each production run in the accelerator increased the generalizability of the models in time. Efforts are now underway to implement a workflow for incorporating human feedback to update the models and evaluating performance on unseen data. The models will eventually be integrated into the existing System Tracking and Reliability system with a web interface for automated anomaly detection and reporting along with a pathway for incorporating human feedback for model updates.

Raj, Anant [ORNL] (ORCID:0000000306711244)

Ultimately Reliable Pyrotechnic Systems

This paper presents the methods by which NASA has designed, built, tested, and certified pyrotechnic devices for high reliability operation in extreme environments and illustrates the potential applications in the oil and gas industry. NASA's extremely successful application of pyrotechnics is built upon documented procedures and test methods that have been maintained and developed since the Apollo Program. Standards are managed and rigorously enforced for performance margins, redundancy, lot sampling, and personnel safety. The pyrotechnics utilized in spacecraft include such devices as small initiators and detonators with the power of a shotgun shell, detonating cord systems for explosive energy transfer across many feet, precision linear shaped charges for breaking structural membranes, and booster charges to actuate valves and pistons. NASA's pyrotechnics program is one of the more successful in the history of Human Spaceflight. No pyrotechnic device developed in accordance with NASA's Human Spaceflight standards has ever failed in flight use. NASA's pyrotechnic initiators work reliably in temperatures as low as -420 F. Each of the 135 Space Shuttle flights fired 102 of these initiators, some setting off multiple pyrotechnic devices, with never a failure. The recent landing on Mars of the Opportunity rover fired 174 of NASA's pyrotechnic initiators to complete the famous '7 minutes of terror.' Even after traveling through extreme radiation and thermal environments on the way to Mars, every one of them worked. These initiators have fired on the surface of Titan. NASA's design controls, procedures, and processes produce the most reliable pyrotechnics in the world. Application of pyrotechnics designed and procured in this manner could enable the energy industry's emergency equipment, such as shutoff valves and deep-sea blowout preventers, to be left in place for years in extreme environments and still be relied upon to function when needed, thus greatly enhancing safety and operational availability.

Scott, John H.

System and Software Reliability (C103)

Within the last decade better reliability models (hardware. software, system) than those currently used have been theorized and developed but not implemented in practice. Previous research on software reliability has shown that while some existing software reliability models are practical, they are no accurate enough. New paradigms of development (e.g. OO) have appeared and associated reliability models have been proposed posed but not investigated. Hardware models have been extensively investigated but not integrated into a system framework. System reliability modeling is the weakest of the three. NASA engineers need better methods and tools to demonstrate that the products meet NASA requirements for reliability measurement. For the new models for the software component of the last decade, there is a great need to bring them into a form that they can be used on software intensive systems. The Statistical Modeling and Estimation of Reliability Functions for Systems (SMERFS'3) tool is an existing vehicle that may be used to incorporate these new modeling advances. Adapting some existing software reliability modeling changes to accommodate major changes in software development technology may also show substantial improvement in prediction accuracy. With some additional research, the next step is to identify and investigate system reliability. System reliability models could then be incorporated in a tool such as SMERFS'3. This tool with better models would greatly add value in assess in GSFC projects.

Wallace, Dolores

A Reliable Earth Return System for Safe Recovery of Mars Samples

The objective of a Mars sample return mission is to bring selected Mars surface materials to Earth. Numerous approaches for the Earth-return segment have been analyzed including propulsive or aerocapture return to low-Earth orbit followed by Space Shuttle rendezvous and direct entry. Of these approaches, ballistic entry of a small capsule terminating in a ground landing has been shown to be the lowest risk strategy. Over the past two years, significant work has been performed towards development of a robust direct entry vehicle for Mars sample return. In June 1999, the NASA Planetary Protection Officer provided initial guidance to the former Mars Sample Return Project. The sample return phase of the mission was assigned a restricted Earth return planetary protection classification. The draft mission requirement states that the total mean probability of release of unsterilized Mars material into the Earth;s biosphere must be less than 1.0E-06 (1 in a million). This strict requirement drives the approach and design of the Earth return system. To meet this requirement, selection of the Earth return strategy and development of the Earth return system must be guided by risk, not performance, based decisions. An initial Probabilistic Risk Assessment (PRA) was performed to address the direct entry Earth return system containment assurance reliability and to identify high-risk elements of this system. The results of this PRA identified risk elements that include thermal protection system performance during entry, spin-eject orientation and aerodynamic stability during entry, structural integrity under atmospheric deceleration and impact loads, and tracking/recovery of this system. This initial probabilistic risk quantification demonstrates that, with the proper development program, a prototypical direct entry design can satisfy the containment assurance reliability requirement. Through the current Mars Sample Return Advanced Technology Development effort, an extensive design, analysis, and test program is presently proceeding with the aim of reducing the containment assurance risk of this system. This technology development effort, guided by a continuing PRA, focuses on key risk areas of a direct entry Earth return system including: the thermal protection system, impact dynamics, structural performance, aerodynamic stability, and ground recovery. This development program will culminate in a system validation flight test, 1-2 years prior to launch of the flight system. This flight test would include the launch, entry, and recovery of a full-scale Earth return system, as a scientific validation of the key risk elements to verify nominal design performance. The results of the initial PRA suggested several dominant failure sequences that can be validated in a flight test. These include: demonstrating the thermal protection system reliability and performance during entry, demonstrating the spin-eject orientation and aero-dynamic stability during entry, demonstrating the structural integrity under atmospheric deceleration and impact loads, and demonstrating tracking and recovery of the Earth return system. This single test will directly address over 50% of the total containment assurance risk elements. This presentation will begin by presenting the relative risk of various Earth return strategies. The results of the initial probabilistic risk assessment will be presented followed by a discussion of the development accomplishments and plans for demonstration of a highly reliable direct entry Earth return system.

Braun, R.

The X-43 Fin Actuation System Problem - Reliability in Shades of Gray

Following the loss of the first X-43 during launch, the mishap investigation board indicated the Fin Actuator System (FAS) needed to have a larger torque margin. To supply this added torque, a second actuator was added. The consequences of what seemed to be a simple modification would trouble the X-43 program. Because of the second actuator, a new computer board was required. This proved to be subject to electronic noise. This resulted in the actuator latch up in ground tests of the FAS for the second launch. Such a latch up would cause the Pegasus booster to fail, as the FAS was a single string system. The problem was corrected and the second flight was successful. The same modifications were added to the FAS for flight three. When the FAS underwent ground tests, it also latched up. The failure indicated that each computer board had a different tolerance to electronic noise. The problem with the FAS was corrected. Subsequently, another failure occurred, raising questions about the design, and the probability of failure for the X-43 Mach 10 flight. This was not simply a technical issue, but illuminated the difficulties facing both managers and engineers in assessing risk, design requirements, and probabilities in cutting edge aerospace projects.

Peebles, Curtis

Calculation of the binomial survivor function

A method is presented for calculating the binomial SF (cumulative binomial distribution), binfc(k;p,n), especially for a large n, beyond the range of existing tables, where conventional computer programs fail because of underflow and overflow, and Gaussian or Poisson approximations yield insufficient accuracy for the purpose at hand. This method is used to calculate and sum the individual binomial terms while using multiplication factors to avoid underflow; the factors are then divided out of the partial sum whenever it has the potential to overflow. A computer program uses this technique to calculate the binomial SF for arbitrary inputs of k, p, and n. Two other algorithms are presented to determine the value of p needed to yield a specified SF for given values of k and n and calculate the value where p = SF for a given k and n. Reliability applications of each algorithm/program are given, e.g., the value of p needed to achieve a stated k-out-of-n:G system reliability and the value of p for which k-out-of-n:G system reliability equals p.

Bowerman, Paul N.

Strapdown platforms using redundant two-degree-of-freedom gyros

A new concept of high reliability strapdown attitude sensing systems for space vehicles is presented. Each system utilizes a set of redundant two-degree-of-freedom gyros. An optimum system configuration is obtained for maximum system reliability and the best measurement accuracy. Improved accuracy of the final data is obtained by using the least-square data reduction technique. Each system possesses a sensor performance management feature which is capable of failure detection, faulty gyro identification, system reconfiguration, and possibly, sensor recalibration. Improvement in reliability as compared to other types of strapdown systems is demonstrated. Details of the development are described in terms of a system containing four gyros.

Hung, J. C.

High-reliability strapdown platforms using two-degree-of-freedom gyros.

This paper presents a new concept of high-reliability strapdown attitude sensing systems for space vehicles. Each system utilizes a set of redundant two-degree-of-freedom gyros. An optimum system configuration is obtained for maximum system reliability and the best measurement accuracy. Improved accuracy of the final data is obtained by using the least-square data reduction technique. Each system possesses a 'sensor performance management' feature which is capable of failure detection, faulty gyro identification, system reconfiguration, and, possibly, sensor recalibration. Improvement in reliability, as compared to other types of strapdown systems, is demonstrated. Details of the development are described in terms of a system containing four gyros.

Hung, J. C.

Field-based AFDD for refrigerant undercharge in residential HVAC systems: enhancing reliability through false alarm mitigation

This study evaluated rule-based and machine learning (ML) based automated fault detection and diagnostics (AFDD) algorithms for detecting refrigerant undercharge faults in residential heating, ventilation, and air conditioning (HVAC) systems, using actual building data and a minimal set of features. The ML-based algorithms included Decision Tree (DT) and K-Nearest Neighbors (KNN). Both the rule-based and ML-based algorithms demonstrated the capability to detect refrigerant undercharge faults of -30% or more. Both types of algorithms exhibited false alarms before the implementation of a false alarm mitigation algorithm, which motivated the development of such a mitigation strategy. After applying the mitigation, false alarms were substantially reduced, with the rule-based algorithm decreasing to 0.6% and the ML-based algorithms reaching 0%, while maintaining strong detection performance. Although the rule-based algorithm initially showed lower performance compared to the ML-based algorithms, its detection accuracy improved after mitigation to a level comparable to the ML-based algorithms. These results confirm that combining false alarm mitigation with both rule-based and ML-based AFDD algorithms significantly enhances practical reliability while preserving robust fault detection capabilities. Furthermore, the findings demonstrate the potential for field deployment of these algorithms in residential HVAC systems and highlight the importance of minimizing false alarms.

False Alarm

DC Microgrid Reliability Enhancement with Adaptive Converter Thermal Management

Due to the different device selections, aging levels, and thermal dissipation performance, some converters may take additional thermal stress on switching devices than others in paralleled converter systems, which will reduce system reliability. To address this problem, this paper proposes a power-sharing strategy with adaptive thermal management. First, the temperature-based power loss model and electrical-thermal model are established. Based on that, a high-accuracy IGBT junction temperature estimate considering the power loss-temperature coupling can be achieved. Further, the thermal-sharing for all the switching devices in paralleled converters can be achieved with the proposed adaptive thermal management strategy. The proposed strategy can change the power-sharing ratio adaptively according to the system operation conditions, which will contribute to the system reliability enhancement. The effectiveness of the proposed strategy is verified through PLECS thermal simulation and joint real-time simulation with Dspace and RT-box.

DC microgrid

DSN Wide Area Network Architecture, Capacity and Performance

This paper discusses the architecture of the wide area network that connects key communications facilities within the National Aeronautic and Space Administration (NASA) Deep Space Network (DSN). Several considerations are given to the design of this wide area network to ensure a timely, reliable, and secure data delivery between the mission users and their spacecraft. The network star configuration simplifies data delivery to users and minimizes operational cost. The dual-path connections maximize the system reliability, with geographical diversity in data routing to avoid single point of failure. Data encryption enhances the protection of mission users’ data. The system bandwidth is determined by balancing the needs to minimize the operating bandwidth cost and to have sufficient bandwidth to be able to deliver data to all users within the required. The DSN uses a class base weighted fair queuing (CBWFQ) method in its data delivery. This scheme guarantees a minimum bandwidth to each class of users and allows users to also access any unused bandwidth by other groups. The paper will also show performance of system reliability and bandwidth margin.

Liao, Jason

Applicability and Limitations of Reliability Allocation Methods

Reliability allocation process may be described as the process of assigning reliability requirements to individual components within a system to attain the specified system reliability. For large systems, the allocation process is often performed at different stages of system design. The allocation process often begins at the conceptual stage. As the system design develops, more information about components and the operating environment becomes available, different allocation methods can be considered. Reliability allocation methods are usually divided into two categories: weighting factors and optimal reliability allocation. When properly applied, these methods can produce reasonable approximations. Reliability allocation techniques have limitations and implied assumptions that need to be understood by system engineers. Applying reliability allocation techniques without understanding their limitations and assumptions can produce unrealistic results. This report addresses weighting factors, optimal reliability allocation techniques, and identifies the applicability and limitations of each reliability allocation technique.

Reliability allocation