Search NASASearch

SEARCH · Search NASA

Results for “Reliability growth”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Study of space shuttle orbiter system management computer function. Volume 1: Analysis, baseline design

A system analysis of the shuttle orbiter baseline system management (SM) computer function is performed. This analysis results in an alternative SM design which is also described. The alternative design exhibits several improvements over the baseline, some of which are increased crew usability, improved flexibility, and improved growth potential. The analysis consists of two parts: an application assessment and an implementation assessment. The former is concerned with the SM user needs and design functional aspects. The latter is concerned with design flexibility, reliability, growth potential, and technical risk. The system analysis is supported by several topical investigations. These include: treatment of false alarms, treatment of off-line items, significant interface parameters, and a design evaluation checklist. An in-depth formulation of techniques, concepts, and guidelines for design of automated performance verification is discussed.

Source record

Heroic Reliability Improvement in Manned Space Systems

System reliability can be significantly improved by a strong continued effort to identify and remove all the causes of actual failures. Newly designed systems often have unexpected high failure rates which can be reduced by successive design improvements until the final operational system has an acceptable failure rate. There are many causes of failures and many ways to remove them. New systems may have poor specifications, design errors, or mistaken operations concepts. Correcting unexpected problems as they occur can produce large early gains in reliability. Improved technology in materials, components, and design approaches can increase reliability. The reliability growth is achieved by repeatedly operating the system until it fails, identifying the failure cause, and fixing the problem. The failure rate reduction that can be obtained depends on the number and the failure rates of the correctable failures. Under the strong assumption that the failure causes can be removed, the decline in overall failure rate can be predicted. If a failure occurs at the rate of lambda per unit time, the expected time before the failure occurs and can be corrected is 1/lambda, the Mean Time Before Failure (MTBF). Finding and fixing a less frequent failure with the rate of lambda/2 per unit time requires twice as long, time of 1/(2 lambda). Cutting the failure rate in half requires doubling the test and redesign time and finding and eliminating the failure causes.Reducing the failure rate significantly requires a heroic reliability improvement effort.

life support

The Role of the U.S. Electric Distribution System in Serving Data Center and Other Large Loads

The rapid expansion of data centers in the United States is reshaping how the electric distribution system must plan for and accommodate large load interconnections. This report evaluates the role of the distribution grid in serving these loads, from small edge facilities to hyperscale campuses. Using national datasets, utility filings, and industry studies, we assess demand growth, reliability requirements, interconnection thresholds, and infrastructure needs at substations and feeders. The analysis highlights the mismatch between fast data center development timelines and slower utility planning and construction cycles, as well as strategies such as phased energization, on-site generation, hosting capacity maps, and structured interconnection frameworks. While focused on data centers, the insights also apply to other high-density loads such as advanced manufacturing, hydrogen production, and electrified transportation. The report concludes with approaches to align planning processes, transparency tools, and regulatory frameworks so utilities can manage new large loads in ways that support a reliable and resilient grid.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Failure rate analysis of Goddard Space Flight Center spacecraft performance during orbital life

Space life performance data on 57 Goddard Space Flight Center spacecraft are analyzed from the standpoint of determining an appropriate reliability model and the associated reliability parameters. Data from published NASA reports, which cover the space performance of GSFC spacecraft launched in the 1960-1970 decade, form the basis of the analyses. The results of the analyses show that the time distribution of 449 malfunctions, of which 248 were classified as failures (not necessarily catastrophic), follow a reliability growth pattern that can be described with either the Duane model or a Weibull distribution. The advantages of both mathematical models are used in order to: identify space failure rates, observe chronological trends, and compare failure rates with those experienced during the prelaunch environmental tests of the flight model spacecraft.

Norris, H. P.

Software reliability modeling and analysis

A discrete and, as approximation to it, a continuous model for the software reliability growth process are examined. The discrete model is based on independent multinomial trials and concerns itself with the joint distribution of the first occurrence time of its underlying events (bugs). The continuous model is based on the order statistics of N independent nonidentically distributed exponential random variables. It is shown that the spacings between bugs are not necessarily independent or exponentially (geometrically) distributed. However, there is a statistical rationale for viewing them so conditionally. Some identifiability problems are pointed out and resolved. In particular, it appears that the number of bugs in a program is not identifiable. Estimated upper bounds and confidence bounds for the residual program eror content are given based on the spacings of the first k bugs removed.

Scholz, F.-W.

Evaluation of competing software reliability predictions

Different software reliability models can produce very different answers when called upon to predict future reliability in a reliability growth context. Users need to know which, if any, of the competing predictions are trustworthy. Some techniques are presented which form the basis of a partial solution to this problem. Rather than attempting to decide which model is generally best, the approach adopted here allows a user to decide upon the most appropriate model for each application.

Abdel-Ghaly, A. A.

An overview of the mathematical and statistical analysis component of RICIS

Mathematical and statistical analysis components of RICIS (Research Institute for Computing and Information Systems) can be used in the following problem areas: (1) quantification and measurement of software reliability; (2) assessment of changes in software reliability over time (reliability growth); (3) analysis of software-failure data; and (4) decision logic for whether to continue or stop testing software. Other areas of interest to NASA/JSC where mathematical and statistical analysis can be successfully employed include: math modeling of physical systems, simulation, statistical data reduction, evaluation methods, optimization, algorithm development, and mathematical methods in signal processing.

Hallum, Cecil R.

Cryogenic optical systems and instruments IV; Proceedings of the Meeting, San Diego, CA, July 10-12, 1990

Consideration is given to cryogenic system design and optical technology; cryogenic instruments, sensors, and detectors; space cryogenic dewars and coolers; and cryogenic mechanisms, testing, and structures. Particular attention is given to mission optimization of the Space Infrared Telescope Facility (SIRTF), alternative aperture stop position designs for SIRTF, scaling laws for lightweight optics, evaluation of a far-infrared Ge:Ga multiplexed detector array, cryogenic limb array etalon spectrometer calibration, reliability growth of coolers for advanced optical systems and instruments, flight-qualified solid argon cooler for the BBXRT instrument, precision mechanisms for optical alignments at cryogenic temperatures, versatile cryogenic rotary-positioning systems, and optical alignments of the Cosmic Background Explorer observatory.

Melugin, Ramsey K.

Statistical modeling of software reliability

This working paper discusses the statistical simulation part of a controlled software development experiment being conducted under the direction of the System Validation Methods Branch, Information Systems Division, NASA Langley Research Center. The experiment uses guidance and control software (GCS) aboard a fictitious planetary landing spacecraft: real-time control software operating on a transient mission. Software execution is simulated to study the statistical aspects of reliability and other failure characteristics of the software during development, testing, and random usage. Quantification of software reliability is a major goal. Various reliability concepts are discussed. Experiments are described for performing simulations and collecting appropriate simulated software performance and failure data. This data is then used to make statistical inferences about the quality of the software development and verification processes as well as inferences about the reliability of software versions and reliability growth under random testing and debugging.

Miller, Douglas R.

The infeasibility of quantifying the reliability of life-critical real-time software

This paper affirms that the quantification of life-critical software reliability is infeasible using statistical methods, whether these methods are applied to standard software or fault-tolerant software. The classical methods of estimating reliability are shown to lead to exorbitant amounts of testing when applied to life-critical software. Reliability growth models are examined and also shown to be incapable of overcoming the need for excessive amounts of testing. The key assumption of software fault tolerance - separately programmed versions fail independently - is shown to be problematic. This assumption cannot be justified by experimentation in the ultrareliability region, and subjective arguments in its favor are not sufficiently strong to justify it as an axiom. Also, the implications of the recent multiversion software experiments support this affirmation.

Butler, Ricky W.

The Infeasibility of Quantifying the Reliability of Life-Critical Real-Time Software

This paper affirms that the quantification of life-critical software reliability is infeasible using statistical methods whether applied to standard software or fault-tolerant software. The classical methods of estimating reliability are shown to lead to exhorbitant amounts of testing when applied to life-critical software. Reliability growth models are examined and also shown to be incapable of overcoming the need for excessive amounts of testing. The key assumption of software fault tolerance separately programmed versions fail independently is shown to be problematic. This assumption cannot be justified by experimentation in the ultrareliability region and subjective arguments in its favor are not sufficiently strong to justify it as an axiom. Also, the implications of the recent multiversion software experiments support this affirmation.

Butler, Ricky W.

An Evolvable Approach to Launch Vehicles for Exploration

This paper presents ideas that may be used individually or in combination to mitigate high costs for separate developments of new crew and heavy-lift cargo launch vehicles, while providing the foundation for a highly reliable and evolvable approach to exploration. Consideration is given to reclassification of cargo for launch purposes into high value versus low value categories, rather than the presently-defined crew versus cargo categories. Objectives for the reclassification are to reduce the gap between payload mass requirements for crew and cargo payloads to better allow closure on a single moderately-sized common core vehicle to reduce development cost, achieve an economical balance between launch frequency and payload mass, and to improve total mission reliability and safety, as compared a light-weight crew vehicle and heavy cargo lift approach. Concepts to reduce design and flight qualification costs for a common core vehicle with derivatives are presented. Appropriate types and mass of cargo for each class of vehicle are identified. Utilization of existing infrastructure and flight hardware is considered to reduce costs and build on proven capabilities. The approach enables low-risk incorporation of international and commercial launch of relatively low-cost, easily replaceable assets as a means to evolve toward longer-duration and more distant missions. Benefits are identified for ground idrastructure, personnel, training, logistics, spares, and system evolution. Technology needs are compared with needs for other aspects of exploration. Technology development phasing, demonstration, and reliability growth opportunities are considered. Flexibility to adapt to future technologies such as advanced in-space propulsion is contrasted with an approach of sizing the cargo launch vehicle based on today's in-space propellants.

Cheuvront, David L.

Shuttle Risk Progression by Flight

Understanding the early mission risk and progression of risk as a vehicle gains insights through flight is important: . a) To the Shuttle Program to understand the impact of re-designs and operational changes on risk. . b) To new programs to understand reliability growth and first flight risk. . Estimation of Shuttle Risk Progression by flight: . a) Uses Shuttle Probabilistic Risk Assessment (SPRA) and current knowledge to calculate early vehicle risk. . b) Shows impact of major Shuttle upgrades. . c) Can be used to understand first flight risk for new programs.

Hamlin, Teri

The Use of Crow-AMSAA Plots to Assess Mishap Trends

Crow-AMSAA (CA) plots are used to model reliability growth. Use of CA plots has expanded into other areas, such as tracking events of interest to management, maintenance problems, and safety mishaps. Safety mishaps can often be successfully modeled using a Poisson probability distribution. CA plots show a Poisson process in log-log space. If the safety mishaps are a stable homogenous Poisson process, a linear fit to the points in a CA plot will have a slope of one. Slopes of greater than one indicate a nonhomogenous Poisson process, with increasing occurrence. Slopes of less than one indicate a nonhomogenous Poisson process, with decreasing occurrence. Changes in slope, known as "cusps," indicate a change in process, which could be an improvement or a degradation. After presenting the CA conceptual framework, examples are given of trending slips, trips and falls, and ergonomic incidents at NASA (from Agency-level data). Crow-AMSAA plotting is a robust tool for trending safety mishaps that can provide insight into safety performance over time.

Dawson, Jeffrey W.

A Multi-Purpose Modular Electronics Integration Node for Exploration Extravehicular Activity

As NASA works to develop an effective integrated portable life support system design for exploration Extravehicular activity (EVA), alternatives to the current system s electrical power and control architecture are needed to support new requirements for flexibility, maintainability, reliability, and reduced mass and volume. Experience with the current Extravehicular Mobility Unit (EMU) has demonstrated that the current architecture, based in a central power supply, monitoring and control unit, with dedicated analog wiring harness connections to active components in the system has a significant impact on system packaging and seriously constrains design flexibility in adapting to component obsolescence and changing system needs over time. An alternative architecture based in the use of a digital data bus offers possible wiring harness and system power savings, but risks significant penalties in component complexity and cost. A hybrid architecture that relies on a set of electronic and power interface nodes serving functional models within the Portable Life Support System (PLSS) is proposed to minimize both packaging and component level penalties. A common interface node hardware design can further reduce penalties by reducing the nonrecurring development costs, making miniaturization more practical, maximizing opportunities for maturation and reliability growth, providing enhanced fault tolerance, and providing stable design interfaces for system components and a central control. Adaptation to varying specific module requirements can be achieved with modest changes in firmware code within the module. A preliminary design effort has developed a common set of hardware interface requirements and functional capabilities for such a node based on anticipated modules comprising an exploration PLSS, and a prototype node has been designed assembled, programmed, and tested. One instance of such a node has been adapted to support testing the swingbed carbon dioxide and humidity control element in NASA s advanced PLSS 2.0 test article. This paper will describe the common interface node design concept, results of the prototype development and test effort, and plans for use in NASA PLSS 2.0 integrated tests.

Hodgson, Edward

Supportability Challenges, Metrics, and Key Decisions for Future Human Spaceflight

Future crewed missions beyond Low Earth Orbit (LEO) represent a logistical challenge that is unprecedented in human space flight. Astronauts will travel farther and stay in space for longer than any previous mission, far from timely abort or resupply from Earth. Under these conditions, supportability { defined as the set of system characteristics that influence the logistics and support required to enable safe and effective operations of systems { will be a much more significant driver of space system lifecycle properties than it has been in the past. This paper presents an overview of supportability for future human space flight. The particular challenges of future missions are discussed, with the differences between past, present, and future missions highlighted. The relationship between supportability metrics and mission cost, performance, schedule, and risk is also discussed. A set of pro- posed strategies for managing supportability is presented (including reliability growth, uncertainty reduction, level of repair, commonality, redundancy, In-Space Manufacturing (ISM) (including the use of material recycling and In-Situ Resource Utilization (ISRU) for spares and maintenance items), reduced complexity, and spares inventory decisions such as the use of predeployed or cached spares - along with a discussion of the potential impacts of each of those strategies. References are provided to various sources that describe these supportability metrics and strategies, as well as associated modeling and optimization techniques, in greater detail. Overall, supportability is an emergent system characteristic and a holistic challenge for future system development. System designers and mission planners must carefully consider and balance the supportability metrics and decisions described in this paper in order to enable safe and effective beyond-LEO human space flight.

Owens, Andrew C.

Field Programmable Gate Array Failure Rate Estimation Guidelines for Launch Vehicle Fault Tree Models

Today's launch vehicles complex electronic and avionics systems heavily utilize Field Programmable Gate Array (FPGA) integrated circuits (IC) for their superb speed and reconfiguration capabilities. Consequently, FPGAs are prevalent ICs in communication protocols such as MILSTD- 1553B and in control signal commands such as in solenoid valve actuations. This paper will identify reliability concerns and high level guidelines to estimate FPGA total failure rates in a launch vehicle application. The paper will discuss hardware, hardware description language, and radiation induced failures. The hardware contribution of the approach accounts for physical failures of the IC. The hardware description language portion will discuss the high level FPGA programming languages and software/code reliability growth. The radiation portion will discuss FPGA susceptibility to space environment radiation.

Al Hassan, Mohammad

Field Programmable Gate Array Failure Rate Estimation Guidelines for Launch Vehicle Fault Tree Models

Today's launch vehicles complex electronic and avionic systems heavily utilize the Field Programmable Gate Array (FPGA) integrated circuit (IC). FPGAs are prevalent ICs in communication protocols such as MIL-STD-1553B, and in control signal commands such as in solenoid/servo valves actuations. This paper will demonstrate guidelines to estimate FPGA failure rates for a launch vehicle, the guidelines will account for hardware, firmware, and radiation induced failures. The hardware contribution of the approach accounts for physical failures of the IC, FPGA memory and clock. The firmware portion will provide guidelines on the high level FPGA programming language and ways to account for software/code reliability growth. The radiation portion will provide guidelines on environment susceptibility as well as guidelines on tailoring other launch vehicle programs historical data to a specific launch vehicle.

Al Hassan, Mohammad