Search NASA⌕ Search

DOE OSTI · 2480022

From Failure to Insight: Analyzing Disk Breakdowns in Large-Scale HPC Environments

Abstract

Disk failure data provides valuable insights for preventing failures, enhancing storage robustness, guiding system design and deployment, and ensuring reliable operations at data centers. This paper introduces two disk failure datasets collected from large-scale HPC production environments over the past five years, comprising over 5,000 failure records from more than 40,000 disks. We analyzed these datasets across multiple dimensions, including temporal, spatial, and relational trends, and performed a comprehensive reliability assessment. Our analysis yielded numerous observations and insights that influence various operational aspects of HPC storage systems. We believe this study offers a holistic understanding of disk failure trends likely to interest the HPC storage community.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

George, Anjus, Wang, Meng, Hanley, Jesse, Ransom, Garrett Wilson, Bent, John, Zimmer, Christopher. 2024-11-01. From Failure to Insight: Analyzing Disk Breakdowns in Large-Scale HPC Environments. https://doi.org/10.1109/scw63240.2024.00070

Cite the original work for its findings. Save a collection to share your selection of sources.