DOE OSTI · 1998847
PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems
Abstract
The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Choudhary, Alok N., Agrawal, Ankit, Liao, Wei-Keng. 2023-05-31. PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems. https://doi.org/10.2172/1998847
Cite the original work for its findings. Save a collection to share your selection of sources.