Search NASA⌕ Search

Engineering topics

Yu, Xiaodong

Publications and source records attributed to Yu, Xiaodong.

2022 AI Testbed Expeditions Report

By exploiting the coherent properties of a light source, coherent diffraction imaging (CDI) is able to obtain the sample image at a nanoscale resolution using the measured diffraction pattern. Bragg Coherent Diffraction Imaging (BCDI) has become valuable for recovering the displacement and strain field of crystals, providing a valuable tool in material science and solid-state physics. X-ray ptychography is another emerging CDI technique that can produce a high-resolution image of the extended sample and has become popular in many research areas (e.g., materials science, biology, electronics, and optics characterization). CDI including BCDI and ptychography has become an established technique in Synchrotron Facilities including the Advanced Photon Source (APS) and will greatly benefit from the 100x coherent flux increase of the upcoming APS Upgrade (APSU). The current image formation process in CDI employs iterative phase retrieval algorithms, which is a time-consuming and computationally expensive process. Especially after APSU, the traditional iterative methods will not be able to match the experimental data acquisition speed. We employ deep learning (DL) approach to replace the iterative approaches, therefore allowing hundreds of times faster recovery of the object. We developed AutoPhaseNN, a DL-based approach which learns to solve the inverse problem without labeled data. Taking 3D BCDI as a representative technique, AutoPhaseNN has been demonstrated to be one hundred times faster than traditional iterative phase retrieval methods while providing comparable image quality. The current network is trained with 64 x 64 x 64 data size, to achieve higher resolution imaging, we will need to scale the network to input and train/infer 3D arrays of size 256 x 256 x 256 (today) and of size 2560x2560x2560 (APSU). However, the scalability of the network is restricted due to the memory-intensive training process. To perform the training for a 256 x 256 x 256 data size, the required memory exceeds the capacity of the current machine. In this project, we explore using Sambanova system to train the network for the direct data inversion for CDI.

36 MATERIALS SCIENCE↗

Scalable and accurate multi-GPU-based image reconstruction of large-scale ptychography data

Abstract While the advances in synchrotron light sources, together with the development of focusing optics and detectors, allow nanoscale ptychographic imaging of materials and biological specimens, the corresponding experiments can yield terabyte-scale volumes of data that can impose a heavy burden on the computing platform. Although graphics processing units (GPUs) provide high performance for such large-scale ptychography datasets, a single GPU is typically insufficient for analysis and reconstruction. Several works have considered leveraging multiple GPUs to accelerate the ptychographic reconstruction. However, most of these works utilize only the Message Passing Interface to handle the communications between GPUs. This approach poses inefficiency for a hardware configuration that has multiple GPUs in a single node, especially while reconstructing a single large projection, since it provides no optimizations to handle the heterogeneous GPU interconnections containing both low-speed (e.g., PCIe) and high-speed links (e.g., NVLink). In this paper, we provide an optimized intranode multi-GPU implementation that can efficiently solve large-scale ptychographic reconstruction problems. We focus on the maximum likelihood reconstruction problem using a conjugate gradient (CG) method for the solution and propose a novel hybrid parallelization model to address the performance bottlenecks in the CG solver. Accordingly, we have developed a tool, called PtyGer ( Pty chographic G PU(multipl e )-based r econstruction), implementing our hybrid parallelization model design. A comprehensive evaluation verifies that PtyGer can fully preserve the original algorithm’s accuracy while achieving outstanding intranode GPU scalability.

97 MATHEMATICS AND COMPUTING↗

Ultrafast Error-bounded Lossy Compression for Scientific Datasets

Today's scientific high-performance computing applications and advanced instruments are producing vast volumes of data across a wide range of domains, which impose a serious burden on data transfer and storage. Error-bounded lossy compression has been developed and widely used in the scientific community because it not only can significantly reduce the data volumes but also can strictly control the data distortion based on the user-specified error bound. Existing lossy compressors, however, cannot offer ultrafast compression speed, which is highly demanded by numerous applications or use cases (such as in-memory compression and online instrument data compression). In this paper, we propose a novel ultrafast error-bounded lossy compressor that can obtain fairly high compression performance on both CPUs and GPUs and with reasonably high compression ratios. The key contributions are threefold. (1) We propose a generic error-bounded lossy compression framework---called SZx---that achieves ultrafast performance through its novel design comprising only lightweight operations such as bitwise and addition/subtraction operations, while still keeping a high compression ratio. (2) We implement SZx on both CPUs and GPUs and optimize the performance according to their architectures. (3) We perform a comprehensive evaluation with six real-world production-level scientific datasets on both CPUs and GPUs. Experiments show that SZx is 2~16x faster than the second-fastest existing error-bounded lossy compressor (either SZ or ZFP) on CPUs and GPUs, with respect to both compression and decompression.

Yu, Xiaodong↗