Search NASA⌕ Search

SEARCH · Search NASA

Results for “Disk failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Alpine disk failure dataset

This dataset consists of 4114 disk failure events collected from Alpine from 07/01/2019 to 04/18/2022. Individual events are identified by their timestamp in ISO 8601 format, rack and enclosure where the failed disk is located.

97 MATHEMATICS AND COMPUTING↗

Collection of Disk Failure Events from Alpine, the Parallel File System for Summit Supercomputer

This dataset contains disk (HDD) failure events collected from the Alpine storage system of the Summit supercomputer, hosted at OLCF, spanning from January 4, 2019, to December 21, 2023 (a total of 4 years, 11 months, and 18 days), covering 89% of its operational lifetime. It includes 3,766 disk failure events, each recorded with its detection timestamp (in ISO 8601 format) and detailed by its location within the storage system - rack, enclosure, and drive slot number.

97 MATHEMATICS AND COMPUTING↗

Disk Failure Dataset from the Campaign Storage System

This dataset consists of 1,389 disk (HDD) failure events collected from the Campaign storage system at LANL. The Campaign system supported various compute platforms throughout its lifespan, including Cielo, Fire, Ice, and notably, the Trinity supercomputer. Each recorded event includes its detection timestamp (in ISO 8601 format) and details such as its location within the storage system—rack, enclosure, and drive slot number. The data, spanning from May 4, 2021, to July 25, 2023 (2 years, 2 months, and 22 days), represents failure events from the terminal years of Campaign's operational period, accounting for 26% of its total operational time.

97 MATHEMATICS AND COMPUTING↗

From Failure to Insight: Analyzing Disk Breakdowns in Large-Scale HPC Environments

Disk failure data provides valuable insights for preventing failures, enhancing storage robustness, guiding system design and deployment, and ensuring reliable operations at data centers. This paper introduces two disk failure datasets collected from large-scale HPC production environments over the past five years, comprising over 5,000 failure records from more than 40,000 disks. We analyzed these datasets across multiple dimensions, including temporal, spatial, and relational trends, and performed a comprehensive reliability assessment. Our analysis yielded numerous observations and insights that influence various operational aspects of HPC storage systems. We believe this study offers a holistic understanding of disk failure trends likely to interest the HPC storage community.

George, Anjus↗

Rotor fragment protection program: Statistics on aircraft gas turbine engine rotor failures that occurred in US commercial aviation during 1979

Statistical information relating to the number of gas turbine engine rotor failures which occurred during 1979 in commercial aviation service use is provided. The predominant failure mode involved blade fragments, 84 percent of which were contained. No uncontained disk failures occurred and although fewer rotor rim and seal failures occurred, 100 percent and 50 percent, respectively, were uncontained. Sixty-eight percent of the 157 rotor failures occurred during the take-off and climb stages of flight.

Delucia, R. A.↗

Fatigue Crack Growth Behavior Evaluation of Grainex Mar-M 247 for NASA's High Temperature, High Speed Turbine Seal Test Rig

The fatigue crack growth behavior of Grainex Mar-M 247 is evaluated for NASA s Turbine Seal Test Facility. The facility is used to test air-to-air seals primarily for use in advanced jet engine applications. Because of extreme seal test conditions of temperature, pressure, and surface speeds, surface cracks may develop over time in the disk bolt holes. An inspection interval is developed to preclude catastrophic disk failure by using experimental fatigue crack growth data. By combining current fatigue crack growth results with previous fatigue strain-life experimental work, an inspection interval is determined for the test disk. The fatigue crack growth life of the NASA disk bolt holes is found to be 367 cycles at a crack depth of 0.501 mm using a factor of 2 on life at maximum operating conditions. Combining this result with previous fatigue strain-life experimental work gives a total fatigue life of 1032 cycles at a crack depth of 0.501 mm. Eddy-current inspections are suggested starting at 665 cycles since eddy current detection thresholds are currently at 0.381 mm. Inspection intervals are recommended every 50 cycles when operated at maximum operating conditions.

Delgado, Irebert R.↗

Integrated Monitoring of Macroalgae Farms Using Acoustics and UUV Sensing

The vision of this project was to develop an integrated system for autonomous underwater vehicle (AUV) monitoring of offshore kelp farms using acoustic, environmental, and optical sensors. This project supports the overall MARINER goals of developing an offshore kelp aquaculture industry to produce low-carbon or carbon-neutral biofuels. The project commenced in the spring of 2018 and used laboratory experiments to test the efficacy of acoustic sensors for monitoring kelp farm lines and growing kelp biomass. Sensors were then integrated onto two AUVs as well as an autonomous surface vehicle in order to establish the optimal type, price point, and vehicle to most efficiently monitor kelp farm structures, kelp biomass, and the surrounding environment. Multiple field deployments of these vehicles and sensors in Massachusetts, New Hampshire, and Maine confirmed that farm lines and kelp could be visualized and that monitoring the spatial patterns of environmental variables was possible. The COVID-19 pandemic severely limited fieldwork activities and laboratory testing, with deployments around kelp farms not occurring again until January 2021. Time spent away from the field focused on data visualization and the development of a low-cost, vessel-based sensor system. Unfortunately, the low-cost system experienced a disk failure during its first deployment and due to multiple resignations from the project team further engineering and development was not possible. Additionally, final results from acoustic sensor testing in 2021 were also not able to be completed due to the resignation of the postdoc leading the analysis. Despite these setbacks, multiple avenues for further development of optical imagery processing from the 360-degree Kelpcam camera and testing of the low-cost sensor system may be possible.

09 BIOMASS FUELS↗

Tutorial: Performance and reliability in redundant disk arrays

A disk array is a collection of physically small magnetic disks that is packaged as a single unit but operates in parallel. Disk arrays capitalize on the availability of small-diameter disks from a price-competitive market to provide the cost, volume, and capacity of current disk systems but many times their performance. Unfortunately, relative to current disk systems, the larger number of components in disk arrays leads to higher rates of failure. To tolerate failures, redundant disk arrays devote a fraction of their capacity to an encoding of their information. This redundant information enables the contents of a failed disk to be recovered from the contents of non-failed disks. The simplest and least expensive encoding for this redundancy, known as N+1 parity is highlighted. In addition to compensating for the higher failure rates of disk arrays, redundancy allows highly reliable secondary storage systems to be built much more cost-effectively than is now achieved in conventional duplicated disks. Disk arrays that combine redundancy with the parallelism of many small-diameter disks are often called Redundant Arrays of Inexpensive Disks (RAID). This combination promises improvements to both the performance and the reliability of secondary storage. For example, IBM's premier disk product, the IBM 3390, is compared to a redundant disk array constructed of 84 IBM 0661 3 1/2-inch disks. The redundant disk array has comparable or superior values for each of the metrics given and appears likely to cost less. In the first section of this tutorial, I explain how disk arrays exploit the emergence of high performance, small magnetic disks to provide cost-effective disk parallelism that combats the access and transfer gap problems. The flexibility of disk-array configurations benefits manufacturer and consumer alike. In contrast, I describe in this tutorial's second half how parallelism, achieved through increasing numbers of components, causes overall failure rates to rise. Redundant disk arrays overcome this threat to data reliability by ensuring that data remains available during and after component failures.

Gibson, Garth A.↗

Redundant disk arrays: Reliable, parallel secondary storage

During the past decade, advances in processor and memory technology have given rise to increases in computational performance that far outstrip increases in the performance of secondary storage technology. Coupled with emerging small-disk technology, disk arrays provide the cost, volume, and capacity of current disk subsystems, by leveraging parallelism, many times their performance. Unfortunately, arrays of small disks may have much higher failure rates than the single large disks they replace. Redundant arrays of inexpensive disks (RAID) use simple redundancy schemes to provide high data reliability. The data encoding, performance, and reliability of redundant disk arrays are investigated. Organizing redundant data into a disk array is treated as a coding problem. Among alternatives examined, codes as simple as parity are shown to effectively correct single, self-identifying disk failures.

Gibson, Garth Alan↗

MLEC-Sim: A Simulator for Evaluating Multi-Level Erasure Coding

We present MLEC-Sim, a sophisticated simulator for Multi-Level Erasure Coding (MLEC), developed in approximately 13 KLOC. The simulator is engineered to analyze the impact of various system configurations and erasure coding policies on system durability and network overhead. It supports a comprehensive range of parameters including disk capacity, disk I/O bandwidth, failure rates, network bandwidth, and system scale, accommodating various erasure coding approaches such as Single-Level Erasure Coding (SLEC), Multi-Level Erasure Coding (MLEC), and Local Reconstruction Codes (LRC). MLEC-Sim provides support for multiple chunk placement policies, including clustered parity and declustered parity, and encompasses a variety of repair methods like Repair-ALL, Repair-FCO, Repair-HYB, and Repair-MIN. It is capable of simulating disk failures through a variety of means, including distribution-based or trace-based mechanisms, and can handle complex multi-level (de)clustered placements and repair processes. A key feature of MLEC-Sim is its adoption of the splitting simulation method for evaluating system durabilities at extremely high levels, which are challenging to assess with traditional simulation approaches. This feature allows for a detailed evaluation of system resilience under a range of conditions, aiding in the selection of appropriate erasure coding solutions for enhancing system durability. MLEC-Sim contributes to the field of data storage and reliability by providing a tool for the detailed evaluation of the durability and efficiency of erasure coding configurations, intended for use by researchers and practitioners in the design and optimization of storage systems.

Wang, Meng↗

Redundant Disk Arrays in Transaction Processing Systems

We address various issues dealing with the use of disk arrays in transaction processing environments. We look at the problem of transaction undo recovery and propose a scheme for using the redundancy in disk arrays to support undo recovery. The scheme uses twin page storage for the parity information in the array. It speeds up transaction processing by eliminating the need for undo logging for most transactions. The use of redundant arrays of distributed disks to provide recovery from disasters as well as temporary site failures and disk crashes is also studied. We investigate the problem of assigning the sites of a distributed storage system to redundant arrays in such a way that a cost of maintaining the redundant parity information is minimized. Heuristic algorithms for solving the site partitioning problem are proposed and their performance is evaluated using simulation. We also develop a heuristic for which an upper bound on the deviation from the optimal solution can be established.

Mourad, Antoine Nagib↗

Braking System for Wind Turbines

Operating turbine stopped smoothly by fail-safe mechanism. Windturbine braking systems improved by system consisting of two large steel-alloy disks mounted on high-speed shaft of gear box, and brakepad assembly mounted on bracket fastened to top of gear box. Lever arms (with brake pads) actuated by spring-powered, pneumatic cylinders connected to these arms. Springs give specific spring-loading constant and exert predetermined load onto brake pads through lever arms. Pneumatic cylinders actuated positively to compress springs and disengage brake pads from disks. During power failure, brakes automatically lock onto disks, producing highly reliable, fail-safe stops. System doubles as stopping brake and "parking" brake.

Krysiak, J. E.↗

The Competition of Failure Modes in an Additively Manufactured Disk Superalloy

Additive manufacturing of powder metallurgy disk superalloys can produce unique microstructures that are different from those usually encountered in traditional processing by consolidation, forging, and heat treatments. Unusual variations in grain size, major and minor phase precipitate sizes, and defects can occur. The associated failure modes of these unique microstructures are of high interest. The objective of this study was to compare the failure modes for a powder metallurgy disk superalloy LSHR produced by electron beam melting additive manufacturing. Specimens were subsequently given different solution heat treatments and a fixed aging heat treatment. Tensile, creep, and fatigue failure modes were screened in tests at elevated temperatures. Failure modes were considered with respect to these unique microstructures.

additive manufacturing↗

Competition of Failure Modes in an Additively Manufactured Disk Superalloy

Additive manufacturing of powder metallurgy (PM) disk superalloys can produce unique microstructures that differ from those usually encountered in traditional processing as the result of consolidation, forging, and heat treatments. Unusual variations in grain size, major and minor phase precipitate sizes, and defects can occur. The failure modes associated with these unique microstructures are of high interest. The objective of this study was to compare the failure modes for a low solvus, high refractory (LSHR) PM disk superalloy produced by electron-beam-melting additive manufacturing. Specimens were subsequently given different solution heat treatments and a fixed aging heat treatment. Tensile, creep, and fatigue failure modes were screened in tests at elevated temperatures. Failure modes were considered with respect to these unique microstructures.

additive manufacturing↗

Advanced turbine disk designs to increase reliability of aircraft engines

Results of analytical studies to improve the low cycle fatigue lives and reliability of turbine disks in high performance gas turbine engines are presented. Advanced disk concepts were evaluated for the first stage high pressure turbines of the CF6-50 and JT8D-17 engines. The advanced disk designs are compared to the existing disks on the bases of cycles to crack initiation and overspeed capability for initially unflawed disks, crack propagation cycles to failure for initially flawed disks, and available kinetic energy of disk burst fragments.

Kaufman, A.↗

Advanced turbine disk designs to increase reliability of aircraft engines

Results of analytical studies to improve the low cycle fatigue lives and reliability of turbine disks in high-performance gas turbine engines are presented. Advanced disk concepts were evaluated for the first-stage high pressure turbines of the CF6-50 and JT8D-17 engines. The advanced disk designs are compared to the existing disks on the bases of cycles to crack initiation and overspeed capability for initially unflawed disks, crack propagation cycles to failure for initially flawed disks, and available kinetic energy of disk burst fragments.

Kaufman, A.↗

Fatigue Life of a NiCr-Coated Powder Metallurgy Disk Superalloy After Varied Processing and Exposures

A protective ductile NiCr coating has shown promise to mitigate oxidation and corrosion attack on superalloy disk alloys. The effects of this coating on fatigue life and failure modes of the disk superalloy are an important concern. The objective of this study was to investigate the fatigue life and failure modes of disk superalloy specimens protected by this coating, using varied pre-coating and post-coating processes. Cylindrical gage fatigue specimens of a powder metallurgy-processed disk superalloy were grit blast or wet blast before being coated with a ductile NiCrY coating, then shot peened at low or medium levels after coating. All were then heat treated, some exposed, and finally all were subjected to fatigue at high temperature. The effects of varied pre-coating treatment, post-coating shot peening, and oxidation plus hot corrosion exposures on fatigue life with the coating were compared.

coating↗

Rotor fragment protection program: Statistics on aircraft gas turbine ngine rotor failures that occurred in U.S. commercial aviation during 1978

This report presents statistical information relating to the number of gas turbine engine rotor failures which occurred in commercial aviation service use. The predominant failure involved blade fragments, 82.4 percent of which were contained. Although fewer rotor rim, disk, and seal failures occurred, 33.3%, 100% and 50% respectively were uncontained. Sixty-five percent of the 166 rotor failures occurred during the takeoff and climb stages of flight.

Delucia, R. A.↗