Search NASA⌕ Search

Engineering topics

Lee, Sunwoo

Publications and source records attributed to Lee, Sunwoo.

Variance-aware weight quantization of multi-level resistive switching devices based on Pt/LaAlO3/SrTiO3 heterostructures

Abstract Resistive switching devices have been regarded as a promising candidate of multi-bit memristors for synaptic applications. The key functionality of the memristors is to realize multiple non-volatile conductance states with high precision. However, the variation of device conductance inevitably causes the state-overlap issue, limiting the number of available states. The insufficient number of states and the resultant inaccurate weight quantization are bottlenecks in developing practical memristors. Herein, we demonstrate a resistive switching device based on Pt/LaAlO 3 /SrTiO 3 (Pt/LAO/STO) heterostructures, which is suitable for multi-level memristive applications. By redistributing the surface oxygen vacancies, we precisely control the tunneling of two-dimensional electron gas (2DEG) through the ultrathin LAO barrier, achieving multiple and tunable conductance states (over 27) in a non-volatile way. To further improve the multi-level switching performance, we propose a variance-aware weight quantization (VAQ) method. Our simulation studies verify that the VAQ effectively reduces the state-overlap issue of the resistive switching device. We also find that the VAQ states can better represent the normal-like data distribution and, thus, significantly improve the computing accuracy of the device. Our results provide valuable insight into developing high-precision multi-bit memristors based on complex oxide heterostructures for neuromorphic applications.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A case study on parallel HDF5 dataset concatenation for high energy physics data analysis

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size, creating a data management challenge for scientists. In order to better facilitate data management, transfer, and analysis on large scale platforms, it is advantageous to aggregate data further into a smaller number of larger files. However, this translation process can consume significant time and resources, and if performed incorrectly the resulting aggregated files can be inefficient for highly parallel access during analysis on large scale platforms. In this paper, we present our case study on parallel I/O strategies and HDF5 features for reducing data aggregation time, making effective use of compression, and ensuring efficient access to the resulting data during analysis at scale. We focus on NOvA detector data in this case study, a large-scale HEP experiment generating many terabytes of data. Here, the lessons learned from our case study inform the handling of similar datasets, thus expanding community knowledge related to this common data management task.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency

Synchronous SGD with data parallelism, the most popular parallelization strategy for CNN training, suffers from the expensive communication cost of averaging gradients among all workers. The iterative parameter updates of SGD cause frequent communications and it becomes the performance bottleneck. In this paper, we propose a lazy parameter update algorithm that adaptively adjusts the parameter update frequency to address the expensive communication cost issue. Our algorithm accumulates the gradients if the difference of the accumulated gradients and the latest gradients is sufficiently small. Here, the less frequent parameter updates reduce the per-iteration communication cost while maintaining the model accuracy. Our experimental results demonstrate that the lazy update method remarkably improves the scalability while maintaining the model accuracy. For ResNet50 training on ImageNet, the proposed algorithm achieves a significantly higher speedup (739.6 on 2048 Cori KNL nodes) as compared to the vanilla synchronous SGD (276.6) while the model accuracy is almost not affected (<0.2% difference).

97 MATHEMATICS AND COMPUTING↗