Search NASA⌕ Search

Engineering topics

Long, Darrell

Publications and source records attributed to Long, Darrell.

Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows

This report details recent progress for the ASCR funded project “Integrated End-to-end Performance Prediction and Diagnosis for Extreme Scientific Workflows”. We refer to the project as IPPD/2, reflecting the 2017 renewal under expanded scope and partners In IPPD/2, we increased our research scope to include data motion. We are focusing on three major aspects: a) observe how data is generated, distributed, and used; b) analyze how data is (repeatedly) consumed with a focus both on repeated patterns and anomalies; and c) explore how to optimize data motion. This new work on data motion will augment and complement IPPD/2’s research that focused on the computational aspects of tasks. We leverage and extend our existing tools and demonstrate our work on the Belle II workflow suite as well as on workflows from NSLS-II. The highlights of our work are as follows: Provenance for Workflows: Provenance is used to provide information enabling quality control, re-run computational workflows, and reproduce results. IPPD/2 has been building a scalable provenance management system that enables the capture of provenance from the high-level workflow through all relevant system levels in one integrated environment. Leveraging this work, our recent efforts have included using provenance as an enabling technique. Workload characterization: Leveraging provenance and analysis, we characterize data movement within network, storage, and memory over a variety of workloads. This characterization enables an understanding by performance analysts and application developers of the range of behaviors that could be expected. Performance Prediction for Workflows: The goal of modeling distributed workflows is to understand performance bottlenecks and enable more intelligent task scheduling to optimize selected metrics of interest (e.g., task throughput or output data rate). IPPD/2 has utilized both analytical and AI/ML modeling methodologies for performance modeling. Advanced Scheduling and Fault Modeling for Workflows: Scheduling of large-scale scientific workflows on geographically distributed resources is a challenging problem. To improve workflow throughput, we combined novel scheduling algorithms with task predictions from performance modeling and fault modeling. Dynamically Alleviating Bottlenecks in Workflows: Exploiting our provenance, analysis, and modeling efforts, we have explored and developed several techniques for dynamically detecting and alleviating bottlenecks in data movement. In particular, we have spent considerable effort demonstrating our techniques on production-like workflow configurations.

97 MATHEMATICS AND COMPUTING↗

WinnowML: Stable feature selection for maximizing prediction accuracy of time-based system modeling

Online deep learning (ODL) has become an important methodology for modeling time-based performance of computer systems. An open problem is the intelligent selection of features from raw workload traces of computer systems. The best methods are overly sensitive to noisy data, causing frequent feature changes and re-training. Using all available features inflates training time and introduces model artifacts if some features should have been dropped. We present WinnowML, a method for automatically determining the most relevant feature subset for a predictive time-series model. WinnowML combines existing feature ranking algorithms and a history of each feature's ranking to iteratively rank a feature set to lower prediction error and maximize long term relevance. From this ranked feature set, the most relevant and stable subset is selected to train a model. Experimentally, we show how WinnowML can lower a model's mean absolute relative error up to 42% on average compared to the closest performing approach. Additionally, we lower the fluctuation in feature ranking and selection up to 65%. We also demonstrate how to combine WinnowML and a model search tool to provide improvements in performance of up to 14.5% when compared to using all the feature available.

Bel, Oceane MS↗

Geomancy: Automated Performance Enhancement through Data Layout Optimization

Large distributed storage systems such as high- performance computing (HPC) systems used by national or international laboratories require sufficient performance and scale for demanding scientific workloads and must handle shifting workloads with ease. Ideally, data is placed in locations to optimize performance, but the size and complexity of large storage systems inhibit rapid effective restructuring of data layouts to maintain performance as workloads shift. To address these issues, we have developed Geomancy, a tool that models the placement of data within a distributed storage system and reacts to drops in performance. Using a combination of machine learning techniques suitable for temporal modeling, Geomancy determines when and where a bottleneck may happen due to changing workloads and suggests changes in the layout that mitigate or prevent them. Our approach to optimizing throughput offers benefits for storage systems such as avoiding potential bottlenecks and increasing overall I/O throughput from 11% to 30%.

Bel, Oceane M.↗