Search NASASearch

SEARCH · Search NASA

Results for “Reliability growth”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Field Programmable Gate Array Reliability Analysis Guidelines for Launch Vehicle Reliability Block Diagrams

Field Programmable Gate Arrays (FPGAs) integrated circuits (IC) are one of the key electronic components in today's sophisticated launch and space vehicle complex avionic systems, largely due to their superb reprogrammable and reconfigurable capabilities combined with relatively low non-recurring engineering costs (NRE) and short design cycle. Consequently, FPGAs are prevalent ICs in communication protocols and control signal commands. This paper will identify reliability concerns and high level guidelines to estimate FPGA total failure rates in a launch vehicle application. The paper will discuss hardware, hardware description language, and radiation induced failures. The hardware contribution of the approach accounts for physical failures of the IC. The hardware description language portion will discuss the high level FPGA programming languages and software/code reliability growth. The radiation portion will discuss FPGA susceptibility to space environment radiation.

Al Hassan, Mohammad

We Can't Count on Repairing All Failures Going to Mars

Reliability analysis often assumes that a complex system can be kept operating indefinitely with scheduled maintenance and emergency repair using a stock of spare parts, as long as the spare parts are not depleted. This assumption seems justified for well-tested, widely used, long operational systems with a multigenerational history of failure, redesign, and reliability growth. It seems doubtful that newer, relatively untried, high technology space systems can always be repaired. We cannot assume space systems will have a low rate of random failures that can all be repaired with a few identical spares. New untried systems usually have a high initial failure rate, called infant mortality, due to errors in requirements, design, parts, materials, and operations planning. These problems can cause groups of related failures called Common Cause Failures (CCFs). The practical definition of a CCF is any failure mode that cannot be cured using identical redundant systems or spare parts. Systems with CCFs may fail repeatedly for the same reason. Can a life support system be kept operating on the way to Mars using only redundant systems and spare parts? The failure history of International Space Station (ISS) life support systems suggests that CCFs are likely to occur and will probably require design changes rather than being reparable with spare parts.

Mars

Assessment of Crew Time for Maintenance and Repairs Activities for Lunar Surface Missions

NASA is currently evaluating different methods to predict how much time crewmembers will spend conducting repair and maintenance activities on future space missions. As mission scope and spacecraft architectures change, it will be necessary to understand how crew repair and maintenance timelines are impacted by mission operations and technology changes. Past work has been done using historical ISS data to accurately predict crew habitation and operation timelines, resulting in the development of NASA’s Exploration Crew Time Model (ECTM). However, understanding crew maintenance and repair requirements has posed a unique challenge due to the complexity of available datasets, the probabilistic nature of sub-system failures, and the impacts of reliability growth on failure rates. This paper presents a methodology to collect and condition empirical repair and maintenance time data from available data sets, to extrapolate from that data to estimate projected maintenance and repair times for a lunar Surface Habitat, and to assess how uncertainty in repair time could impact utilization time on the lunar surface. NASA International Space Station (ISS) maintenance and crew time data are logged into two central databases, the Maintenance Data Collection (MDC) and the Operations Planning Timeline Integration System (OPTimIS) respectively. Separately, each of these two datasets capture only portions of the complete set of data required to generate an accurate assessment of crew time spent on maintenance activities at a sub-system level. MDC provides a detailed catalog of failure events and an overview of the failure’s required maintenance and OPTimIS provides a description of crew activities and crew time durations dedicated to maintenance. To create a more useful crew time estimate for maintenance timelines, the authors developed a methodology to capture relevant data from each set and combine and utilize that data by linking crew time requirements to specific components. The authors compare the failure logs in the MDC to crew activity logs pulled from OPTimIS and then process the data to estimate required repair times for each failure event. Data is also classified by the outcome of each repair event, whether the failed component was replaced or whether it was repaired in place. The entire maintenance activity dataset is then categorized based on the class of failed component to allow for a statistically significant sample size for each class and to provide accurate crew time estimates for any components lacking relevant data. This resultant component repair time data can be used in the future to generate Mean Time To Repair (MTTR) estimates and confidence intervals for each class of component based on a probabilistic distribution of documented maintenance events. These improved MTTR values can then be applied to candidate element sub-system architectures, along with component Mean Time Between Failure (MTBF) data to generate distributions for potential required system crew repair time estimates for a given mission. Repair time distributions can then be used to develop more accurate crew schedules and to assess potential available utilization time.

Crew Time

Assessment of Crew Time for Maintenance and Repair Activities for Lunar Surface Missions

NASA is currently evaluating different methods to predict how much time crewmembers will spend conducting repair and maintenance activities on future space missions. As mission scope and spacecraft architectures change, understanding how crew repair and maintenance timelines are impacted by mission operations and technology changes is vital for future mission planning. Past work has been done using historical International Space Station (ISS) data to accurately predict crew habitation and operation timelines, resulting in the development of NASA’s Exploration Crew Time Model (ECTM). However, understanding crew maintenance and repair requirements has posed a unique challenge due to the complexity of available datasets, the probabilistic nature of sub-system failures, and the impacts of reliability growth on failure rates. This paper presents a methodology to collect and condition empirical repair and maintenance time data from available datasets, to extrapolate from that data to estimate projected maintenance and repair times for a lunar Surface Habitat (SH), and to assess how uncertainty in repair time could impact utilization time on the lunar surface. NASA ISS maintenance and crew time data are logged into two central databases: the Maintenance Data Collection (MDC) and the Operations Planning Timeline Integration System (OPTimIS). Separately, each of these two datasets capture only portions of the complete set of data required to generate an accurate assessment of crew time spent on maintenance activities at a sub-system level. To create a more useful crew time estimate for maintenance timelines, the authors developed a methodology to capture relevant data from each set and combine and utilize that data by linking crew time requirements to specific components. The authors compare the failure logs in the MDC to crew activity logs pulled from OPTimIS and then process the data to estimate required repair time for each failure and repair event. The entire maintenance activity dataset is then categorized based on the class of failed component to ensure a significant sample size for each class and accurate crew time estimates for any components lacking relevant data. This resultant component repair time data can be used in the future to generate Mean Time to Repair (MTTR) estimates and confidence intervals for each class of component based on a probabilistic distribution of documented maintenance events. These improved MTTR values can then be applied to candidate element sub-system architectures, along with component Mean Time Between Failure (MTBF) data to generate distributions for potential required system crew repair time estimates for a given mission. The authors applied these modeling methods to a case study of a crewed mission to the planned SH and produced expected corrective maintenance crew time distributions. The results produced an expected corrective maintenance crew time at over 24 hours per mission, and a maintenance crew time distribution that reflects the importance of planning for sufficient maintenance requirements each mission. Repair time distributions can then be used to develop more accurate crew schedules and to assess potential available utilization time.

Crew Time

The Growth of Mass and Reliability Source Database and Me

Mass and Reliability Source (MaRS) is a database which draws together information from multiple other sources and databases (MADS, VMDB, ISS PART, Bayesian PowerPoint, etc.). Information such as mass, operating hours, and failure data are compiled to create a tool for the building of probabilistic Risk Assessment (PRA) models. MaRS also has great potential for deep space mission planning. Understanding the lowest point of failure within an ORU helps to inform the efficient selection of spare parts necessary for a journey where mass and volume are at a premium and resupply is not possible.

George, Cory A.

Highly accelerated life testing (HALT): A review from a statistical perspective

Despite its use in one form or another for at least four decades, HALT and related techniques [e.g., highly accelerated-stress screening (HASS) and stress audits (HASA)] are not well understood within the statistical community and remain controversial. This largely reflects a conflict in motivation between engineers, testing under harsh conditions to discover and eliminate failure modes, and statisticians, taking a more cautious approach to develop quantitative estimates of parameters such as mean time between failures (MTBF). Here, this review article will clarify HALT concepts and methods and explain where it fits within the universe of methods that involve the application of accelerating factors to compress the time required to evaluate or enhance product reliability. A major distinction is between methods such as HALT, a high-stress test-analyze-fix-test iterative process directed at improving reliability by discovering and fixing weak points in a design, and quantitative accelerated life testing (QALT), whose goal is the estimation of product life for a fixed design. We discuss methods such as physics of failure that offer some hope of bridging the gap between the qualitative nature of HALT, and purely quantitative statistical methods. We present a variety of engineering applications of HALT including metal fatigue, piping and pressure vessels, structural damage, radiation damage, and rotating machinery. We also discuss potential synergies between HALT and QALT, such as rapid identification, through HALT, of failure modes requiring quantitative analysis. For further study, extensive references to the applicable literature are provided as well as an appendix that describes related methods.

97 MATHEMATICS AND COMPUTING

Adoption of AI in the Utility T&D Sector: Use Cases, Consequence, Assessment and Benefits

Digital transformation and utilization of artificial intelligence (AI) in the electric grid are fundamentally changing the industry’s approach to common problems and enabling a broader paradigm shift in grid planning and operations. The change in approach is circularly both enabling and driving modernization, with load growth and reliable management of data center and AI infrastructure shifting away from planning approaches with relatively predictable behaviors and toward a mix of consumer and industrial choices that surpass human cognitive abilities to process. This movement has potential to condition humans to not understand the system on which the AI depends, while requiring it for development of the necessary infrastructure. Approaches which would address most likely grid conditions and events, such as faults, aging of equipment, and weather, now must also account for large loads which shift not based upon weather or time of day, but the computational load. Quantifying computational load is independent of the traditional grid forecasting variables, where a data center’s aggregate load is determined by user and AI system behavior and decoupled from normal grid planning and operations. AI is both the cause and solution for these challenges, with new grid planning tools integrating massive amounts of decisions into frameworks.

24 POWER TRANSMISSION AND DISTRIBUTION

Benefits and Challenges of California Offshore Wind Electricity: An Updated Assessment

Offshore wind (OSW) technology has recently been included in California’s plans to achieve 100% carbon-free electricity by 2045. As an emerging technology, many features of OSW are changing more rapidly than established renewable options and are shaped by local circumstances in unique ways that limit transferrable experiences globally. This paper fills a gap in the literature by providing an updated technological assessment of OSW in California to determine its viability and competitiveness in the state’s electricity generation mix to achieve its near-term energy and environmental goals. Through a critical synthesis and extrapolation of technical, social, and economic analyses, we identify several major improvements in its potential. First, we note that while estimates of OSW’s costs per MWh of installed capacity have generally documented and projected a long-term decline, recent technical, microeconomic, and macroeconomic factors have caused significant backsliding of this momentum. Second, we project that the potential dollar value benefits of OSW’s greenhouse gas reduction capabilities have increased by one to two orders of magnitude, primarily due to major upward revisions of the social cost of carbon. Several co-benefits, including enhanced reliability, economic growth, and environmental justice, look to be increasingly promising due to a combination of technological advances and policy initiatives. Despite these advancements, OSW continues to face several engineering and broader challenges. We assess the current status of these challenges, as well as current and future strategies to address them. We conclude that OSW is now overall an even more attractive electricity-generating option than at the beginning of this decade.

Rose, Adam (ORCID:0000000333477684)

Requirements for implementing real-time control functional modules on a hierarchical parallel pipelined system

Analysis of a robot control system leads to a broad range of processing requirements. One fundamental requirement of a robot control system is the necessity of a microcomputer system in order to provide sufficient processing capability.The use of multiple processors in a parallel architecture is beneficial for a number of reasons, including better cost performance, modular growth, increased reliability through replication, and flexibility for testing alternate control strategies via different partitioning. A survey of the progression from low level control synchronizing primitives to higher level communication tools is presented. The system communication and control mechanisms of existing robot control systems are compared to the hierarchical control model. The impact of this design methodology on the current robot control systems is explored.

Wheatley, Thomas E.

Discontinuous pore fluid distribution under microgravity--KC-135 flight investigations

Designing a reliable plant growth system for crop production in space requires the understanding of pore fluid distribution in porous media under microgravity. The objective of this experimental investigation, which was conducted aboard NASA KC-135 reduced gravity flight, is to study possible particle separation and the distribution of discontinuous wetting fluid in porous media under microgravity. KC-135 aircraft provided gravity conditions of 1, 1.8, and 10(-2) g. Glass beads of a known size distribution were used as porous media; and Hexadecane, a petroleum compound immiscible with and lighter than water, was used as wetting fluid at residual saturation. Nitrogen freezer was used to solidify the discontinuous Hexadecane ganglia in glass beads to preserve the ganglia size changes during different gravity conditions, so that the blob-size distributions (BSDs) could be measured after flight. It was concluded from this study that microgravity has little effect on the size distribution of pore fluid blobs corresponding to residual saturation of wetting fluids in porous media. The blobs showed no noticeable breakup or coalescence during microgravity. However, based on the increase in bulk volume of samples due to particle separation under microgravity, groups of particles, within which pore fluid blobs were encapsulated, appeared to have rearranged themselves under microgravity.

manned

GPS/Optical/Inertial Integration for 3D Navigation Using Multi-Copter Platforms

In concert with the continued advancement of a UAS traffic management system (UTM), the proposed uses of autonomous unmanned aerial systems (UAS) have become more prevalent in both the public and private sectors. To facilitate this anticipated growth, a reliable three-dimensional (3D) positioning, navigation, and mapping (PNM) capability will be required to enable operation of these platforms in challenging environments where global navigation satellite systems (GNSS) may not be available continuously. Especially, when the platform's mission requires maneuvering through different and difficult environments like outdoor opensky, outdoor under foliage, outdoor-urban and indoor, and may include transitions between these environments. There may not be a single method to solve the PNM problem for all environments. The research presented in this paper is a subset of a broader research effort, described in [1]. The research is focused on combining data from dissimilar sensor technologies to create an integrated navigation and mapping method that can enable reliable operation in both an outdoor and structured indoor environment. The integrated navigation and mapping design is utilizes a Global Positioning System (GPS) receiver, an Inertial Measurement Unit (IMU), a monocular digital camera, and three short to medium range laser scanners. This paper describes specifically the techniques necessary to effectively integrate the monocular camera data within the established mechanization. To evaluate the developed algorithms a hexacopter was built, equipped with the discussed sensors, and both hand-carried and flown through representative environments. This paper highlights the effect that the monocular camera has on the aforementioned sensor integration scheme's reliability, accuracy and availability.

Dill, Evan T.

High Reliability at Minimum Cost

This paper investigates the minimum cost of improving the reliability of complex technical systems. The two major methods to improve reliability are redesigning the system for higher reliability or providing redundant components to replace failed elements. The costs of redesign for reliability or adding redundancy are estimated. The most cost-effective combination for high reliability can be identified. The cost of increasing the intrinsic reliability of a system can be modeled as cost proportional to 1/(system failure rate) a , where the exponent “a” measures the difficulty of increasing reliability. The “a” exponent can vary from 0.25 to about 2.5. Operational reliability can also be increased by using redundant systems. The failure rate for N parallel redundant units is (system failure rate) N . The cost of redundancy is N times the system cost. The total redundant system cost is proportional to N/(system failure rate) a . The cost of redundancy increases as N gets larger, but larger N allows a higher system failure rate, which reduces the system design cost. There is a certain N, a certain level of redundancy, that has the minimum cost to achieve the required overall redundant system failure rate. The minimum cost for the redundant system is achieved at the optimum level of redundancy. The N for minimum cost is equal to -a ln (redundant system failure rate). The minimum cost of the N redundant systems is proportional to N * (original system failure rate) a . The optimum redesigned individual system failure rate is proportional to exp (-1/a), so the greater the difficulty, the higher the optimum individual system failure rate. Increasing the intrinsic reliability of a system encounters diminishing returns and at some point it becomes more cost-effective to add redundancy. The difficulty of increasing intrinsic system reliability determines the optimum design for high reliability at minimum cost.

reliability

High Reliability at Minimum Cost

This paper investigates the minimum cost of improving the reliability of complex technical systems. The two major methods to improve reliability are redesigning the system for higher reliability or providing redundant components to replace failed elements. The costs of redesign for reliability or adding redundancy are estimated. The most cost-effective combination for high reliability can be identified. The cost of increasing the intrinsic reliability of a system can be modeled as cost proportional to 1/(system failure rate) a , where the exponent “a” measures the difficulty of increasing reliability. The “a” exponent can vary from 0.25 to about 2.5. Operational reliability can also be increased by using redundant systems. The failure rate for N parallel redundant units is (system failure rate) N . The cost of redundancy is N times the system cost. The total redundant system cost is proportional to N/(system failure rate) a . The cost of redundancy increases as N gets larger, but larger N allows a higher system failure rate, which reduces the system design cost. There is a certain N, a certain level of redundancy, that has the minimum cost to achieve the required overall redundant system failure rate. The minimum cost for the redundant system is achieved at the optimum level of redundancy. The N for minimum cost is equal to -a ln (redundant system failure rate). The minimum cost of the N redundant systems is proportional to N * (original system failure rate) a . The optimum redesigned individual system failure rate is proportional to exp (-1/a), so the greater the difficulty, the higher the optimum individual system failure rate. Increasing the intrinsic reliability of a system encounters diminishing returns and at some point it becomes more cost-effective to add redundancy. The difficulty of increasing intrinsic system reliability determines the optimum design for high reliability at minimum cost.

reliability

NASA trend analysis procedures

This publication is primarily intended for use by NASA personnel engaged in managing or implementing trend analysis programs. 'Trend analysis' refers to the observation of current activity in the context of the past in order to infer the expected level of future activity. NASA trend analysis was divided into 5 categories: problem, performance, supportability, programmatic, and reliability. Problem trend analysis uncovers multiple occurrences of historical hardware or software problems or failures in order to focus future corrective action. Performance trend analysis observes changing levels of real-time or historical flight vehicle performance parameters such as temperatures, pressures, and flow rates as compared to specification or 'safe' limits. Supportability trend analysis assesses the adequacy of the spaceflight logistics system; example indicators are repair-turn-around time and parts stockage levels. Programmatic trend analysis uses quantitative indicators to evaluate the 'health' of NASA programs of all types. Finally, reliability trend analysis attempts to evaluate the growth of system reliability based on a decreasing rate of occurrence of hardware problems over time. Procedures for conducting all five types of trend analysis are provided in this publication, prepared through the joint efforts of the NASA Trend Analysis Working Group.

Source record

Poor reliability of public charging stations can impede the growth of the electric vehicle market

How does the reliability of public charging infrastructure affect electric vehicle (EV) adoption? Substantial public and private investments are expanding EV charging networks, but concerns are growing about the poor reliability of existing chargers and its potential impacts on EV adoption. Using data from a nationwide survey, we employ a choice model to quantify the effects of perceived charging reliability on Americans’ intentions to purchase new or used EVs. By randomly assigning participants to receive information characterizing public charging as either very reliable or very unreliable, we show a causal effect of reliability perceptions on EV purchase intentions. In conclusion, we find that differences in perceived reliability are equivalent to changing price by 32 % of purchasing budget or changing range by 366 miles, underscoring the importance of reliable public charging.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Digital Assurance for Grid Reliability in the Era of Large Load Growth

The rapid expansion of large electric loads is reshaping the operational and regulatory landscape of the U.S. electric grid. These facilities are reaching new scales of expansion, now exceeding a gigawatt per site, and their highly sensitive, digitally driven behaviors introduce new reliability risks. Recent grid events, including large load losses following routine transmission disturbances, highlight the consequences of limited ride-through capability, inconsistent protection settings, inadequate modeling, and lack of behind-the-meter visibility. Parallels to earlier integration challenges of new grid technologies suggest that the grid’s existing processes, standards, and interconnection frameworks are no longer adequate for emerging large loads. This brief synthesizes lessons from the evolution of inverter-based resource regulation and applies them to large-load integration. It identifies critical gaps in modeling accuracy, interconnection processes, performance standards, and compliance mechanisms. Technical recommendations emphasize advanced monitoring, improved modeling, coordinated communication protocols, modernized substations, and structured behind-the-meter control schemes. Collectively, these measures provide a roadmap to maintain bulk power system reliability while enabling the continued growth of large, electrified digital infrastructure.

24 - POWER TRANSMISSION AND DISTRIBUTION

Rate limits in silicon sheet growth - The connections between vertical and horizontal methods

Meniscus-defined techniques for the growth of thin silicon sheets fall into two categories: vertical and horizontal growth. The interactions of the temperature field and the crystal shape are analyzed for both methods using two-dimensional finite-element models which include heat transfer and capillarity. Heat transfer in vertical growth systems is dominated by conduction in the melt and the crystal, with almost flat melt/crystal interfaces that are perpendicular to the direction of growth. The high axial temperature gradients characteristic of vertical growth lead to high thermal stresses. The maximum growth rate is also limited by capillarity which can restrict the conduction of heat from the melt into the crystal. In horizontal growth the melt/crystal interface stretches across the surface of the melt pool many times the crystal thickness, and low growth rates are achievable with careful temperature control. With a moderate axial temperature gradient in the sheet a substantial portion of the latent heat conducts along the sheet and the surface of the melt pool becomes supercooled, leading to dendritic growth. The thermal supercooling is surpressed by lowering the axial gradient in the crystal; this configuration is the most desirable for the growth of high quality crystals. An expression derived from scaling analysis relating the growth rate and the crucible temperature is shown to be reliable for horizontal growth.

Thomas, Paul D.