Search NASA⌕ Search

SEARCH · Search NASA

Results for “Reliability analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Gateway Program Safety and Mission Assurance Integration - the Future of Safe Deep Space Human Exploration

As a foundational element of the National Aeronautics and Space Administration (NASA) Artemis Campaign, the Gateway is an incrementally built cislunar spacecraft that will serve as a platform for deep space human exploration, science, and technology demonstration. The Gateway will be a unifying catalyst for international partners around the world to establish sustained deep space scientific investigations, lunar surface access, and missions to Mars. As human exploration moves farther away from Earth, spacecraft designs must prioritize and optimize mass and volume allocations, while minimizing human and spacecraft risk. To accomplish this objective, the Gateway Program Safety and Mission Assurance functions develop, implement, and ensure compliance with requirements, in concert with the accurate characterization and transparent communication of residual hazard risks, for integrated safety, reliability and maintainability and quality assurance. Safety and Mission Assurance was a key contributor during Gateway program pre-formulation and formulation activities where safety and reliability analysis was embedded in the Gateway Systems Engineering and Integration team. During these early program stages, a preliminary Gateway Integrated Hazard Analysis and Preliminary Gateway Probabilistic Risk Assessment assisted in Gateway architectural and operational definition as part of a risk-informed design process. As the deep space architecture has matured, the integrated Safety and Mission Assurance analyses have matured, new safety review processes have been developed, and requirements have been refined to ensure compliance with integrated safety and mission assurance objectives. The Gateway Program is currently concluding the preliminary design review informed milestone, where the primary objectives included: - Ensured completeness and consistency of the preliminary design, including the meeting of all requirements within appropriate margins and acceptable risk posture. - Identification of any major issues moving forward to the Critical Design phase. At this milestone, Safety and Mission Assurance provided numerous products, including Gateway Top Risks and Risk Mitigation Plans, updated integrated hazard analyses, updated probabilistic risk assessment, Crew Survival Analysis Report, and updated Safety and Mission Assurance Requirements and Plans. These products provide a many-faceted perspective on the inherent risk and available mitigations involved in flying the current proposed vehicle design and anticipated stack configurations. In addition, Safety and Mission Assurance identified top technical, process and workforce concerns to be addressed as the program progresses toward the critical design phase. This paper will detail the evolution of the Gateway Program Safety and Mission Assurance integration functions, provide its current status and lessons learned for future human spaceflight programs. Throughout this paper the key tenets of the Gateway Program Safety and Mission Assurance will be discussed: - Application of a risk-informed approach to identify and mitigate areas of highest risk. - Leverage of valuable processes and lessons learned from earlier spaceflight programs. - Development of Safety and Mission Assurance products to inform design risk trades. - Utilization of common Safety and Mission Assurance practices to identify safety risks for multiple perspectives: top-down, bottom-up, and across lines of integration. - Approval of safety hazards at the appropriate level of authority, keeping most deliberation closest to design expertise and elevating risks of greatest concern for program-level consideration. - Championing of Safety and Mission Assurance processes and forums to foster a pervasive safety culture that is transparent, inclusive, and collaborative between all partners. These tenets have allowed the Gateway Safety and Mission Assurance function to play a key role in optimized vehicle design evolution, and early identification and mitigation of Gateway program and Artemis mission risk.

Helen Vaccaro↗

Rancor Integrated Procedure System (RIPS): A Computer-Based Procedure Platform for Advanced Reactor Research

The Rancor Microworld Simulator is a simplified, pressurized water, small modular reactor simulator that includes a multi-unit plant model server, an advanced digital human-machine control interface, and the Rancor Integrated Procedure System (RIPS). Rancor provides a research and development tool that can be used for collecting operator performance data and for prototyping concepts of operations (ConOps) for advanced reactor development. RIPS is meant as a research tool and includes many unique features: (1) RIPS has a robust procedure authoring system. (2) RIPS has the capability to run any of the three IEEE-Std-1786 computer-based procedure types. (3) RIPS can be configured to take on the look and feel of different vendors’ computer-based procedure systems for the purpose of developing and evaluating different ConOps for plant upgrades or new builds. (4) RIPS includes the capability for logging operator procedure use, including integrating procedure logs with Rancor simulator logs, thereby allowing automated data collection of operator scenario runs. (5) RIPS integrates with the Human Unimodel for Nuclear Technology to Enhance Reliability (HUNTER), a dynamic human reliability analysis environment that creates a digital human twin or virtual operator to mimic reactor operator performance. (6) RIPS includes support for automation of plant monitoring and control functions. While RIPS is explicitly built into Rancor, it may also be used with full-scope training simulators. This functionality allows RIPS to be used for existing plants and advanced reactors under development.

99 - GENERAL AND MISCELLANEOUS↗

PSA 2025 DPRA for Cyber Optimization

Cyberattacks can have many different attack paths, durations, and goals. There are also many different mitigation options involving hardware, software, and/or humans. Evaluating defense options should include quantitative evaluation of overall effectiveness to make cost and risk-informed decisions. Typical cyberattack modeling methods only provide a qualitative evaluation and have difficulty with time dependent scenarios. The main areas of cybersecurity are confidentiality, integrity, and availability. For companies with cyber-physical systems such as advanced nuclear reactors, cyber-related integrity is a requirement set by the U.S. Nuclear Regulatory Commission. But companies are also concerned about availability or reliability as a business case. As cyber threats are evolving to a business-for-hire structure, more attacks focus on disrupting business success and reliability, causing financial and economic stability risk. Companies want reliability analysis while optimizing cost, which requires more than safety modeling methods. Dynamic-state-based and Markov-based modeling provides a method for better cyber scenario modeling with timing and conditional features not found in other numerical evaluation methods. EMRALD (Event Modeling Risk Assessment using Lined Diagrams) is a dynamic risk analysis modeling and simulation tool and has features that reduce modeling issues such as state-base explosion found in Markov-based tools. It has been used to model different time-dependent events including plant behavior and operator procedures. As a general modeling tool, EMRALD can also be used to model cyberattack scenarios with varying mitigation options and quantify effectiveness, producing numerical data for risk-informed decisions. This paper uses EMRALD to demonstrate that dynamic risk analysis can be used for cyber threat modeling to provide insights for design decision-making and optimize defense strategies.

97 - MATHEMATICS AND COMPUTING↗

Method of Testing and Predicting Failures of Electronic Mechanical Systems

A method employing a knowledge base of human expertise comprising a reliability model analysis implemented for diagnostic routines is disclosed. The reliability analysis comprises digraph models that determine target events created by hardware failures human actions, and other factors affecting the system operation. The reliability analysis contains a wealth of human expertise information that is used to build automatic diagnostic routines and which provides a knowledge base that can be used to solve other artificial intelligence problems.

Iverson, David L.↗

The application of emulation techniques in the analysis of highly reliable, guidance and control computer systems

Emulation techniques can be a solution to a difficulty that arises in the analysis of the reliability of guidance and control computer systems for future commercial aircraft. Described here is the difficulty, the lack of credibility of reliability estimates obtained by analytical modeling techniques. The difficulty is an unavoidable consequence of the following: (1) a reliability requirement so demanding as to make system evaluation by use testing infeasible; (2) a complex system design technique, fault tolerance; (3) system reliability dominated by errors due to flaws in the system definition; and (4) elaborate analytical modeling techniques whose precision outputs are quite sensitive to errors of approximation in their input data. Use of emulation techniques for pseudo-testing systems to evaluate bounds on the parameter values needed for the analytical techniques is then discussed. Finally several examples of the application of emulation techniques are described.

Migneault, Gerard E.↗

Reliability and Maintainability Analysis of a High Air Pressure Compressor Facility

This paper discusses a Reliability, Availability, and Maintainability (RAM) independent assessment conducted to support the refurbishment of the Compressor Station at the NASA Langley Research Center (LaRC). The paper discusses the methodologies used by the assessment team to derive the repair by replacement (RR) strategies to improve the reliability and availability of the Compressor Station (Ref.1). This includes a RAPTOR simulation model that was used to generate the statistical data analysis needed to derive a 15-year investment plan to support the refurbishment of the facility. To summarize, study results clearly indicate that the air compressors are well past their design life. The major failures of Compressors indicate that significant latent failure causes are present. Given the occurrence of these high-cost failures following compressor overhauls, future major failures should be anticipated if compressors are not replaced. Given the results from the RR analysis, the study team recommended a compressor replacement strategy. Based on the data analysis, the RR strategy will lead to sustainable operations through significant improvements in reliability, availability, and the probability of meeting the air demand with acceptable investment cost that should translate, in the long run, into major cost savings. For example, the probability of meeting air demand improved from 79.7 percent for the Base Case to 97.3 percent. Expressed in terms of a reduction in the probability of failing to meet demand (1 in 5 days to 1 in 37 days), the improvement is about 700 percent. Similarly, compressor replacement improved the operational availability of the facility from 97.5 percent to 99.8 percent. Expressed in terms of a reduction in system unavailability (1 in 40 to 1 in 500), the improvement is better than 1000 percent (an order of magnitude improvement). It is worthy to note that the methodologies, tools, and techniques used in the LaRC study can be used to evaluate similar high value equipment components and facilities. Also, lessons learned in data collection and maintenance practices derived from the observations, findings, and recommendations of the study are extremely important in the evaluation and sustainment of new compressor facilities.

Safie, Fayssal M.↗

Constellation Ground Systems Launch Availability Analysis: Enhancing Highly Reliable Launch Systems Design

Success of the Constellation Program's lunar architecture requires successfully launching two vehicles, Ares I/Orion and Ares V/Altair, within a very limited time period. The reliability and maintainability of flight vehicles and ground systems must deliver a high probability of successfully launching the second vehicle in order to avoid wasting the on-orbit asset launched by the first vehicle. The Ground Operations Project determined which ground subsystems had the potential to affect the probability of the second launch and allocated quantitative availability requirements to these subsystems. The Ground Operations Project also developed a methodology to estimate subsystem reliability, availability, and maintainability to ensure that ground subsystems complied with allocated launch availability and maintainability requirements. The verification analysis developed quantitative estimates of subsystem availability based on design documentation, testing results, and other information. Where appropriate, actual performance history was used to calculate failure rates for legacy subsystems or comparative components that will support Constellation. The results of the verification analysis will be used to assess compliance with requirements and to highlight design or performance shortcomings for further decision making. This case study will discuss the subsystem requirements allocation process, describe the ground systems methodology for completing quantitative reliability, availability, and maintainability analysis, and present findings and observation based on analysis leading to the Ground Operations Project Preliminary Design Review milestone.

Gernand, Jeffrey L.↗

Constellation Ground Systems Launch Availability Analysis: Enhancing Highly Reliable Launch Systems Design

Success of the Constellation Program's lunar architecture requires successfully launching two vehicles, Ares I/Orion and Ares V/Altair, in a very limited time period. The reliability and maintainability of flight vehicles and ground systems must deliver a high probability of successfully launching the second vehicle in order to avoid wasting the on-orbit asset launched by the first vehicle. The Ground Operations Project determined which ground subsystems had the potential to affect the probability of the second launch and allocated quantitative availability requirements to these subsystems. The Ground Operations Project also developed a methodology to estimate subsystem reliability, availability and maintainability to ensure that ground subsystems complied with allocated launch availability and maintainability requirements. The verification analysis developed quantitative estimates of subsystem availability based on design documentation; testing results, and other information. Where appropriate, actual performance history was used for legacy subsystems or comparative components that will support Constellation. The results of the verification analysis will be used to verify compliance with requirements and to highlight design or performance shortcomings for further decision-making. This case study will discuss the subsystem requirements allocation process, describe the ground systems methodology for completing quantitative reliability, availability and maintainability analysis, and present findings and observation based on analysis leading to the Ground Systems Preliminary Design Review milestone.

Gernand, Jeffrey L.↗

Estimation and enhancement of real-time software reliability through mutation analysis

A simulation-based technique for obtaining numerical estimates of the reliability of N-version, real-time software is presented. An extended stochastic Petri net is employed to represent the synchronization structure of N versions of the software, where dependencies among versions are modeled through correlated sampling of module execution times. Test results utilizing specifications for NASA's planetary lander control software indicate that mutation-based testing could hold greater potential for enhancing reliability than the desirable but perhaps unachievable goal of independence among N versions.

Geist, Robert↗

Software reliability modeling and analysis

A discrete and, as approximation to it, a continuous model for the software reliability growth process are examined. The discrete model is based on independent multinomial trials and concerns itself with the joint distribution of the first occurrence time of its underlying events (bugs). The continuous model is based on the order statistics of N independent nonidentically distributed exponential random variables. It is shown that the spacings between bugs are not necessarily independent or exponentially (geometrically) distributed. However, there is a statistical rationale for viewing them so conditionally. Some identifiability problems are pointed out and resolved. In particular, it appears that the number of bugs in a program is not identifiable. Estimated upper bounds and confidence bounds for the residual program eror content are given based on the spacings of the first k bugs removed.

Scholz, F.-W.↗

[MaRS Project]

The Space Exploration Division of the Safety and Mission Assurances Directorate is responsible for reducing the risk to Human Space Flight Programs by providing system safety, reliability, and risk analysis. The Risk & Reliability Analysis branch plays a part in this by utilizing Probabilistic Risk Assessment (PRA) and Reliability and Maintainability (R&M) tools to identify possible types of failure and effective solutions. A continuous effort of this branch is MaRS, or Mass and Reliability System, a tool that was the focus of this internship. Future long duration space missions will have to find a balance between the mass and reliability of their spare parts. They will be unable take spares of everything and will have to determine what is most likely to require maintenance and spares. Currently there is no database that combines mass and reliability data of low level space-grade components. MaRS aims to be the first database to do this. The data in MaRS will be based on the hardware flown on the International Space Stations (ISS). The components on the ISS have a long history and are well documented, making them the perfect source. Currently, MaRS is a functioning excel workbook database; the backend is complete and only requires optimization. MaRS has been populated with all the assemblies and their components that are used on the ISS; the failures of these components are updated regularly. This project was a continuation on the efforts of previous intern groups. Once complete, R&M engineers working on future space flight missions will be able to quickly access failure and mass data on assemblies and components, allowing them to make important decisions and tradeoffs.

Aruljothi, Arunvenkatesh↗

A reliability and comparative analysis of two standby system configurations.

Equations are derived which enable one to calculate the system reliability for parallel or triple modular redundant systems with standby spares. Software error detection is introduced into the TMR/Spares system configuration in order to utilize fully all of the units. An indication of the sensitivity of the system reliability to an increase in the number of spares, partitioning, switching, variations in the powered and unpowered failures rates, and time is presented. A comparison of the parallel and the TMR/Spares system configurations, under similar conditions, is given.

Taylor, D. S.↗

A Framework for Reliability and Safety Analysis of Complex Space Missions

Long duration and complex mission scenarios are characteristics of NASA's human exploration of Mars, and will provide unprecedented challenges. Systems reliability and safety will become increasingly demanding and management of uncertainty will be increasingly important. NASA's current pioneering strategy recognizes and relies upon assurance of crew and asset safety. In this regard, flexibility to develop and innovate in the emergence of new design environments and methodologies, encompassing modeling of complex systems, is essential to meet the challenges.

Safety Analysis↗

Ceramic component reliability with the restructured NASA/CARES computer program

The Ceramics Analysis and Reliability Evaluation of Structures (CARES) integrated design program on statistical fast fracture reliability and monolithic ceramic components is enhanced to include the use of a neutral data base, two-dimensional modeling, and variable problem size. The data base allows for the efficient transfer of element stresses, temperatures, and volumes/areas from the finite element output to the reliability analysis program. Elements are divided to insure a direct correspondence between the subelements and the Gaussian integration points. Two-dimensional modeling is accomplished by assessing the volume flaw reliability with shell elements. To demonstrate the improvements in the algorithm, example problems are selected from a round-robin conducted by WELFEP (WEakest Link failure probability prediction by Finite Element Postprocessors).

Powers, Lynn M.↗

Ceramic component reliability with the restructured NASA/CARES computer program

The Ceramics Analysis and Reliability Evaluation of Structures (CARES) integrated design program on statistical fast fracture reliability and monolithic ceramic components is enhanced to include the use of a neutral data base, two-dimensional modeling, and variable problem size. The data base allows for the efficient transfer of element stresses, temperatures, and volumes/areas from the finite element output to the reliability analysis program. Elements are divided to insure a direct correspondence between the subelements and the Gaussian integration points. Two-dimensional modeling is accomplished by assessing the volume flaw reliability with shell elements. To demonstrate the improvements in the algorithm, example problems are selected from a round-robin conducted by WELFEP (WEakest Link failure probability prediction by Finite Element Postprocessors).

Powers, Lynn M.↗

Accounting for Point Estimate Uncertainty in Space Systems Reliability and Risk Analysis

Understanding and accounting for uncertainty in risk analysis is a critical step in the management and communication of risk in engineered systems. The component and system-level analysis to determine the probability of a negative outcome and its consequence is often quantified by a point estimate. Many Program and Enterprise decisions involving technical concerns and issues rely on reliability engineering activities to produce quantified risk analysis to inform the decision making process. At NASA, it is common to use a Probabilistic Risk Analysis (PRA) to inform the overall risk to Loss of Mission or Loss of Crew that involves integration across all spacecraft subsystem fault trees to produce an overall probability of mission failure. The point estimate is an estimate of this overall probability and is an immediate result of a fault tree model. It is the result of a model where the probability of each event is taken to be equal to its mean. The value provides an approximation of the overall mean without running any uncertainty calculations (e.g., no sampling). Using only the point estimate can lead to a false sense of precision and the point estimate may not match the resulting mean when uncertainty is taken into consideration. This paper will explore five conditions that can cause the PRA model mean to diverge from the point estimate and will provide engineers and managers insight into the importance of understanding uncertainty in the elements of PRA models.

Paul J Collier↗