Search NASA⌕ Search

Engineering topics

Kang, Qiao

Publications and source records attributed to Kang, Qiao.

Data Storage for HEP Experiments in the Era of High-Performance Computing

As particle physics experiments push their limits on both the energy and the intensity frontiers, the amount and complexity of the produced data are also expected to increase accordingly. With such large data volumes, next-generation efforts like the HL-LHC and DUNE will rely even more on both high-throughput (HTC) and high-performance (HPC) computing clusters. Full utilization of HPC resources requires scalable and efficient data-handling and I/O. For the last few decades, ROOT has been used by most HEP experiments to store data. However, other storage technologies like HDF5 may perform better in HPC environments. Initial explorations with HDF5 have begun using ATLAS, CMS and DUNE data; the DUNE experiment has also adopted HDF5 for its data-acquisition system. This paper presents the future outlook of the HEP computing and the role of HPC, and a summary of ongoing and future works to use HDF5 as a possible data storage technology for the HEP experiments to use in HPC environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency

Synchronous SGD with data parallelism, the most popular parallelization strategy for CNN training, suffers from the expensive communication cost of averaging gradients among all workers. The iterative parameter updates of SGD cause frequent communications and it becomes the performance bottleneck. In this paper, we propose a lazy parameter update algorithm that adaptively adjusts the parameter update frequency to address the expensive communication cost issue. Our algorithm accumulates the gradients if the difference of the accumulated gradients and the latest gradients is sufficiently small. Here, the less frequent parameter updates reduce the per-iteration communication cost while maintaining the model accuracy. Our experimental results demonstrate that the lazy update method remarkably improves the scalability while maintaining the model accuracy. For ResNet50 training on ImageNet, the proposed algorithm achieves a significantly higher speedup (739.6 on 2048 Cori KNL nodes) as compared to the vanilla synchronous SGD (276.6) while the model accuracy is almost not affected (<0.2% difference).

97 MATHEMATICS AND COMPUTING↗