Search NASA⌕ Search

SEARCH · Search NASA

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Assessment of Crew Time for Maintenance and Repairs Activities for Lunar Surface Missions

NASA is currently evaluating different methods to predict how much time crewmembers will spend conducting repair and maintenance activities on future space missions. As mission scope and spacecraft architectures change, it will be necessary to understand how crew repair and maintenance timelines are impacted by mission operations and technology changes. Past work has been done using historical ISS data to accurately predict crew habitation and operation timelines, resulting in the development of NASA’s Exploration Crew Time Model (ECTM). However, understanding crew maintenance and repair requirements has posed a unique challenge due to the complexity of available datasets, the probabilistic nature of sub-system failures, and the impacts of reliability growth on failure rates. This paper presents a methodology to collect and condition empirical repair and maintenance time data from available data sets, to extrapolate from that data to estimate projected maintenance and repair times for a lunar Surface Habitat, and to assess how uncertainty in repair time could impact utilization time on the lunar surface. NASA International Space Station (ISS) maintenance and crew time data are logged into two central databases, the Maintenance Data Collection (MDC) and the Operations Planning Timeline Integration System (OPTimIS) respectively. Separately, each of these two datasets capture only portions of the complete set of data required to generate an accurate assessment of crew time spent on maintenance activities at a sub-system level. MDC provides a detailed catalog of failure events and an overview of the failure’s required maintenance and OPTimIS provides a description of crew activities and crew time durations dedicated to maintenance. To create a more useful crew time estimate for maintenance timelines, the authors developed a methodology to capture relevant data from each set and combine and utilize that data by linking crew time requirements to specific components. The authors compare the failure logs in the MDC to crew activity logs pulled from OPTimIS and then process the data to estimate required repair times for each failure event. Data is also classified by the outcome of each repair event, whether the failed component was replaced or whether it was repaired in place. The entire maintenance activity dataset is then categorized based on the class of failed component to allow for a statistically significant sample size for each class and to provide accurate crew time estimates for any components lacking relevant data. This resultant component repair time data can be used in the future to generate Mean Time To Repair (MTTR) estimates and confidence intervals for each class of component based on a probabilistic distribution of documented maintenance events. These improved MTTR values can then be applied to candidate element sub-system architectures, along with component Mean Time Between Failure (MTBF) data to generate distributions for potential required system crew repair time estimates for a given mission. Repair time distributions can then be used to develop more accurate crew schedules and to assess potential available utilization time.

Crew Time↗

Assessment of Crew Time for Maintenance and Repair Activities for Lunar Surface Missions

NASA is currently evaluating different methods to predict how much time crewmembers will spend conducting repair and maintenance activities on future space missions. As mission scope and spacecraft architectures change, understanding how crew repair and maintenance timelines are impacted by mission operations and technology changes is vital for future mission planning. Past work has been done using historical International Space Station (ISS) data to accurately predict crew habitation and operation timelines, resulting in the development of NASA’s Exploration Crew Time Model (ECTM). However, understanding crew maintenance and repair requirements has posed a unique challenge due to the complexity of available datasets, the probabilistic nature of sub-system failures, and the impacts of reliability growth on failure rates. This paper presents a methodology to collect and condition empirical repair and maintenance time data from available datasets, to extrapolate from that data to estimate projected maintenance and repair times for a lunar Surface Habitat (SH), and to assess how uncertainty in repair time could impact utilization time on the lunar surface. NASA ISS maintenance and crew time data are logged into two central databases: the Maintenance Data Collection (MDC) and the Operations Planning Timeline Integration System (OPTimIS). Separately, each of these two datasets capture only portions of the complete set of data required to generate an accurate assessment of crew time spent on maintenance activities at a sub-system level. To create a more useful crew time estimate for maintenance timelines, the authors developed a methodology to capture relevant data from each set and combine and utilize that data by linking crew time requirements to specific components. The authors compare the failure logs in the MDC to crew activity logs pulled from OPTimIS and then process the data to estimate required repair time for each failure and repair event. The entire maintenance activity dataset is then categorized based on the class of failed component to ensure a significant sample size for each class and accurate crew time estimates for any components lacking relevant data. This resultant component repair time data can be used in the future to generate Mean Time to Repair (MTTR) estimates and confidence intervals for each class of component based on a probabilistic distribution of documented maintenance events. These improved MTTR values can then be applied to candidate element sub-system architectures, along with component Mean Time Between Failure (MTBF) data to generate distributions for potential required system crew repair time estimates for a given mission. The authors applied these modeling methods to a case study of a crewed mission to the planned SH and produced expected corrective maintenance crew time distributions. The results produced an expected corrective maintenance crew time at over 24 hours per mission, and a maintenance crew time distribution that reflects the importance of planning for sufficient maintenance requirements each mission. Repair time distributions can then be used to develop more accurate crew schedules and to assess potential available utilization time.

Crew Time↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

Size of metallic and polyethylene debris particles in failed cemented total hip replacements

Reports of differing failure rates of total hip prostheses made of various metals prompted us to measure the size of metallic and polyethylene particulate debris around failed cemented arthroplasties. We used an isolation method, in which metallic debris was extracted from the tissues, and a non-isolation method of routine preparation for light and electron microscopy. Specimens were taken from 30 cases in which the femoral component was of titanium alloy (10), cobalt-chrome alloy (10), or stainless steel (10). The mean size of metallic particles with the isolation method was 0.8 to 1.0 microns by 1.5 to 1.8 microns. The non-isolation method gave a significantly smaller mean size of 0.3 to 0.4 microns by 0.6 to 0.7 microns. For each technique the particle sizes of the three metals were similar. The mean size of polyethylene particles was 2 to 4 microns by 8 to 13 microns. They were larger in tissue retrieved from failed titanium-alloy implants than from cobalt-chrome and stainless-steel implants. Our results suggest that factors other than the size of the metal particles, such as the constituents of the alloy, and the amount and speed of generation of debris, may be more important in the failure of hip replacements.

Non-NASA Center↗

Low Activity Waste Glass Optimization with Property Models from Machine Learning, Part 2: Experimental Validation and Active Learning

The United States Department of Energy is responsible for managing legacy nuclear waste stored in underground tanks at the Hanford Site. To treat the waste, it is planned as the current baseline to separately vitrify low-activity waste (LAW) and high-level waste fractions. Previously, machine learning (ML) based glass property models (e.g., chemical durability, viscosity, electrical conductivity and SO3 solubility) were developed with prediction uncertainties. A waste glass optimization approach was then established to enable the capability of using these ML models in LAW glass formulation. In this study, the previous ML models were first experimentally validated, and the results were incorporated back into the database to update the ML models. The updated models and formulations showed increased waste loading while reducing the failure rate, demonstrating improved predictive accuracy, reduced uncertainties, and the effectiveness of active learning in guiding high-dimensional, nonlinear LAW glass design. This represents the first experimental validation of ML based LAW glass formulation, with practical benefits such as higher waste loading, shorter mission duration, and lower operational risk.

Lu, Xiaonan (ORCID:0000000179708148)↗

Reliability Analysis of Power Grids Considering Component Failures of Variable Energy Resources

This paper proposes an improved model for the reliability assessment of power systems considering component failures of variable energy resources (VER). The inherent intermittency of VER such as solar photovoltaic (PV) and wind farms, along with their susceptibility to component failures, present significant challenges to reliable system operation. These issues, combined with power grid operation and network constraints, complicate the reliable operation of VER-integrated power systems. Here, to address these concerns, this paper introduces a reliability assessment framework that considers VER input variability, its impact on component availability, and their resulting impact on overall system reliability. Stochastic models based on discrete Markov processes are developed to incorporate variable irradiance, wind speeds, and their effects on PV and wind component failure rates. A next-event and state transition-based approach is then developed to integrate the stochastic models into a mixed-timing sequential Monte Carlo simulation framework for composite reliability assessment. Case studies on the RTS-GMLC system demonstrate the effectiveness of the proposed model in evaluating the reliability of VER-integrated systems.

Pandit, Dilip [Sandia National Laboratories (SNL-N↗

Underwater unexploded ordnance discrimination based on intrinsic target polarizabilities – A case study

Seabed unexploded ordnance that resulted partly from the high failure rate among munitions from more than 80 years ago and from decades of military training and testing of weapons systems poses an increasing concern all around the world. Although existing magnetic systems can detect clusters of debris, they are not able to tell whether a munition is still intact requiring special removal (e.g. in situ detonation) or is harmless scrap metal. The marine environment poses unique challenges, and transferring knowledge and approaches from land to a marine environment has not been easy and straightforward. On land, the background soil conductivity is much lower than the conductivity of the unexploded ordnance and the electromagnetic response of a target is essentially the same as that in free space. For those frequencies required for target characterization in the marine environment, the seawater response must be accounted for and removed from the measurements. The system developed for this study uses fields from three orthogonal transmitters to illuminate the target and four three-component receivers to measure the signal arranged in a configuration that inherently cancels the system's response due to the enclosing seawater, the sea–bottom interface and the air–sea interface for shallow deployments. The system was tested as a cued system on land and underwater in San Francisco Bay – it was mounted on a simple platform on top of a support structure that extended 1 m below and allowed the diver to place metal objects to a specific location even in low-visibility conditions. The measurements were stable and repeatable. Furthermore, target responses estimated from marine measurements matched those from land acquisition, confirming that the seawater and air–sea interface responses were removed successfully. Thirty-six channels of normalized induction responses were used for the classification, which was done by estimating the target principal dipole polarizabilities. Our results demonstrated that the system can resolve the intrinsic polarizabilities of the target, with clear distinctions between those of symmetric intact unexploded ordnance and irregular scrap metal. The prototype system was able to classify an object based on its size, shape and metal content and correctly estimate its location and orientation.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

Instrumentation for the Investigation of Pitch Bearing Design and Reliability

Recently, there has been an increasing level of industry interest in pitch system and pitch bearing reliability. Pitch bearings are used in wind turbines to connect the blade root to the hub. Some populations of pitch bearings have demonstrated a 12% failure rate in 20 years. As rotor diameters continue to increase for tall land-based and offshore wind turbines, pitch bearings are becoming even larger in diameter, which can make them vulnerable to deflections and consequent stress concentrations. There is an increased need to more accurately study pitch bearing deformations, misalignment, load distributions, and contact stresses. A significant body of work has investigated fatigue lives and wear characteristics of pitch bearings on ground-based test rigs. NREL has also recently begun a research program related to pitch bearing reliability, recognizing its growing importance for wind turbines. The purpose of this paper is to describe a set of instrumentation that was recently installed on a 1.5 MW wind turbine at the NREL Flatirons Campus and provide an example data set. To the authors' knowledge, this will be the first publicly available pitch bearing data collection campaign on an operational wind turbine.

17 WIND ENERGY↗

Increasing Reliability and Safety of Hydrogen Components - Reliability Data Collection

Come learn about the new Hydrogen Component Reliability Database (HyCReD) and participate in discussions on hydrogen component reliability data collection, collaboration, and analysis. Funded by the U.S. Department of Energy's Office of Energy Efficiency and Renewable Energy under the Hydrogen and Fuel Cell Technologies Office, HyCReD is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety reliability for hydrogen facilities by integrating risk reduction methodologies and component reliability data taxonomies that support hydrogen infrastructure failure rate analysis.

component↗

Performance Test of Mini LVDT - ELVIS

The Institute for Energy (IFE) Technology has been a pioneer in the development of Linear Variable Differential Transformers (LVDTs) for in-pile testing, deploying over 2,200 units in various reactor environments with less than a 10% failure rate after five years of operation. This report focuses on the performance testing of IFE’s Mini LVDT, a compact sensor ideal for material test reactor experiments. The Mini LVDT, with a limited range of +/- 1.5 mm, offers excellent performance comparable to larger LVDTs, making it valuable in space-constrained applications. The development of an Enhanced Linear Variable Intrinsic Sensor (ELVIS) with internal temperature monitoring capabilities represents a significant advancement, addressing the critical need for real-time, accurate measurements in high-radiation and high-temperature environments. Two ELVIS prototypes were evaluated in terms of both temperature and displacement, showcasing promising results, though challenges with noise during temperature measurements were identified. This report summarizes the rigorous testing performed at Idaho National Laboratory (INL) and highlights the potential applications of Mini LVDTs and ELVIS in nuclear and other high-precision industries.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Evaluation of LLM-Generated Kokkos Code Using Compile-Time and Run-Time Testing

Due to the growing use of large language models (LLMs) by developers and researchers, it has become essential to reliably evaluate their ability to generate code that uses specialized libraries. We explore the use of compile-time and run-time evaluation of LLM-generated Kokkos code through extending the methods used by OpenAI with the HumanEval dataset. Our evaluation framework is based on the first 40 prompts from the Kokkos138 dataset. We start by discussing two different forms of LLM prompting, using entirely plain English or providing pseudocode for added context. These two methods are used to generate Kokkos code with the Llama-3.1-8B-Instruct and CodeQwen1.5-7B-Chat models. We found that both forms of prompting led to high failure rates and difficulties with reliably parsing LLM-generated code, while prompts with pseudocode for context generally led to improved results on more complicated tests.

97 MATHEMATICS AND COMPUTING↗

Extending Component Lifetime And Improving Inverter Reliability (ECLAIIR)

Inverter reliability remains one of the most persistent challenges limiting the performance, availability, and economic viability of utility‑scale photovoltaic (PV) plants. Industry data consistently show that inverters account for the highest share of corrective maintenance events and unplanned outages across PV fleets. These failures result in energy losses, increased O&M costs, and reduced confidence in long‑term solar asset performance. Motivated by these challenges, this project—Extending Component Lifetime and Improving Inverter Reliability (ECLAIIR)—was undertaken to systematically investigate inverter degradation and failure mechanisms, develop predictive maintenance capabilities, and establish data‑driven pathways to improve service life and reduce the Levelized Cost of Energy (LCOE) for large‑scale PV systems. The primary goal of the project was to identify pre‑failure signatures in string inverters using both lab‑based accelerated lifetime testing and field‑based data and to develop predictive maintenance algorithms that can anticipate inverter faults before they occur. Through collaboration with inverter testing laboratory, solar PV plant owner, and failure‑analysis experts, the project advanced the technical understanding of inverter reliability. By instrumenting inverters with thermistors, humidity sensors, power‑quality meters, and acoustic sensors, the research established how multiple sensing modalities can reliably detect deviations from normal behavior hours to days before failure. These findings substantially enhance scientific understanding of inverter failure kinetics and provide the PV industry with the most comprehensive cross‑OEM characterization of early‑stage failure indicators reported to date. Technically, the project demonstrated the effectiveness of predictive maintenance by developing and validating the PreDICT (Predictive Diagnostics of PV Inverters Using Condition Monitoring and Trend Analysis) framework—a multi‑layer diagnostic architecture combining peer‑to‑peer analytics, historical trend modeling, and advanced machine‑learning techniques such as the Sequential Conditional Variational Autoencoder (SCVAE). This predictive model achieved more than 90% accuracy in detecting pre‑failure conditions and provided up to four days of lead time before inverter failure in field scenarios. Economically, the project’s LCOE analysis showed that predictive maintenance can reduce lifetime energy losses and minimize corrective maintenance interventions. Modeling indicated that, depending on inverter failure rates and replacement timelines, predictive maintenance can significantly reduce LCOE impacts associated with inverter downtime: from as high as 19.4% under conventional maintenance strategies to 0.1%–10.17% when predictive analytics are adopted. These results confirm that predictive maintenance is both technically feasible and economically advantageous for utilities and plant operators. The project’s findings also have broad public benefit. By improving inverter reliability and reducing downtime, predictive maintenance directly increases electricity generation from existing PV assets. Enhanced reliability lowers operational costs for utilities, which can translate over time into lower energy costs for consumers. Furthermore, the project’s technical publications, conference presentations, and industry workshops ensure that knowledge gained is shared broadly across the solar industry, supporting workforce development and enabling utilities of all sizes to adopt modern asset‑health monitoring practices. The retrofitting case study and service‑life prediction framework further support informed decision‑making for aging PV fleets, helping operators extend system life and reduce electronic waste. In summary, the ECLAIIR project significantly advanced the state of knowledge on inverter degradation, demonstrated the technical and economic value of predictive maintenance, and delivered actionable tools and insights that support more reliable, cost‑effective, and sustainable PV plant operation. The outcomes of this project will continue to inform utility practices, guide inverter design improvements, and strengthen the long‑term performance of solar assets nationwide.

14 SOLAR ENERGY↗

AGR-5/6/7 Irradiation Experiment Fission Product Mass Balance

This report presents the fission product mass balance for the AGR-5/6/7 TRISO fuel irradiation experiment. The fission product inventories deposited on capsule components outside of the fuel (e.g., stainless-steel shells, graphite holders, Grafoil disks, and associated hardware) were quantified as part of the post-irradiation examination (PIE) to assess the performance of this fuel. Comparisons were made between these inventories and depletion calculations, non-destructive measurements of fuel fission product inventories in the intact fuel compacts, prior AGR experiments, and results from among each of the five distinct AGR-5/6/7 capsules. The data served as estimates of the condensable fission products released from the fuel during irradiation. Excluding Capsule 1 (which experienced accidental damage during irradiation) and Capsule 3 (which was tested at very high irradiation temperatures of >1300°C), the results indicate that the AGR-5/6/7 fuel performed comparably to fuel from earlier AGR experiments. The mass balance results were also used to estimate the number of particles with in-pile SiC failures. Subject to the assumptions made in these estimates, and excluding Capsules 1 and 3, the in-pile SiC failure rates for AGR-5/6/7 are comparable to those observed in AGR-2.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Investigation of the Effect of Gate Oxide Screening with Adjustment Pulse on Commercial SiC Power MOSFETs

This paper presents a method to recover the negative threshold voltage shift during high field gate oxide screening of 1.2 kV 4H-SiC MOSFETs with an additional adjustment gate voltage pulse. To reduce field failure rates of the MOSFETs in operation, manufacturers perform a screening treatment to remove devices with extrinsic defects in the oxide. Current gate oxide screening procedures are limited to oxide fields at or below ~9 MV/cm for short durations (<1 s), which is not enough to remove all the devices with extrinsic defects. The results show that by implementing a lower field gate pulse, the threshold voltage shift can be partially recovered, and therefore the maximum screening field and time can be increased. However, both the initial screening pulse and the adjustment pulse require careful calibration to prevent significant degradation of the device threshold voltage, on-resistance, interface state density, or intrinsic lifetime. With a well calibrated set of pulses, higher screening fields can be utilized without significantly damaging the devices. This leads to an improvement in the overall screening efficiency of the process, reducing the number of devices with extrinsic oxide defects entering the field, and improving the reliability of the SiC MOSFETs in operation.

42 ENGINEERING↗

Evaluation of cloud height, optical thickness, and phase retrievals from the CHROMA algorithm applied to Sentinel-3 OLCI data

We previously developed the Cloud Height Retrieval from O 2 Molecular Absorption (CHROMA) algorithm for the Ocean Color Instrument (OCI) on the new NASA Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission. Here, we apply CHROMA to observations from the Ocean Land Colour Instrument (OLCI) to guide expectations for PACE, as it will take some time to obtain large-scale validation data for OCI. We use cloud top height (CTH), phase, and (for liquid clouds) cloud optical thickness (COT) data from the ground-based Atmospheric Radiation Measurement (ARM) network to evaluate the OLCI retrievals. We found that OLCI and Moderate Resolution Imaging Spectroradiometer (MODIS) CTH compare similarly well to the ARM reference. OLCI has a tendency to underestimate CTH as CTH increases, and algorithm assumptions about cloud geometric thickness may contribute to this. ARM COT from multifilter shadowband radiometers (MFRSR) and Sun photometers are well-correlated with one another, albeit with a roughly 30 % offset on average; OLCI and MODIS COT agree more closely with the MFRSR data. OLCI retrieval uncertainty estimates show skill at telling low-uncertainty cases from high-uncertainty ones, although CTH uncertainties are underestimated. Additionally, we compare the OLCI data to satellite retrievals based on thermal infrared measurements from MODIS and Sea and Land Surface Temperature Radiometer (SLSTR) data. Differences are broadly consistent with physical expectations based on the A-band vs. thermal techniques, although one key challenge in such aggregated comparisons is different cloud masking sensitivities and algorithm failure rates meaning additional sampling differences are introduced. We conclude by discussing the transition to and possible enhancements for PACE OCI.

Sayer, Andrew M. [Univ. of Maryland Baltimore Coun↗

Development of reliability prediction technique for semiconductor diodes

New fundamental technique of reliability prediction for semiconductor diodes based on realistic mathematical models can be applied to component failure rate prediction including mechanical degradation, electrical degradation, environmental stress factors, and electrical load stress factors.

Ryerson, C. M.↗