Search NASA⌕ Search

Engineering topics

Yun, Daqing

Publications and source records attributed to Yun, Daqing.

Exploratory analysis and performance prediction of big data transfer in High-performance Networks

Big data transfer in large-scale scientific and business applications is increasingly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) via advance bandwidth reservation. Provisioning agents need to carefully schedule data transfer requests, compute network paths, and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized, could be simply wasted due to the exclusive access during the approved time window, and cause extra overhead and complexity for resource management. This calls for accurate performance prediction to reserve bandwidths that match actual needs and avoid over-provisioning. We employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements collected in the past several years from data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated testbeds. We first analyze the performance patterns in response to a comprehensive list of parameters in end-host systems, network connections, and data transfer applications, which motivate the use of machine learning and also help us identify the effects of latent factors. We then propose threshold- and clustering-based methods to eliminate negative effects of latent factors in data preprocessing and build a robust performance predictor based on customized domain-oriented loss functions. The performance of the proposed methods is verified by extensive experiments using SVR and RFR as well as theoretical analysis of the general performance bound.

97 MATHEMATICS AND COMPUTING↗

On Performance Prediction of Big Data Transfer in High-performance Networks

Big data generated by large-scale scientific and industrial applications need to be transferred between different geographical locations for remote storage, processing, and analysis. High-speed dedicated connections provisioned in High-performance Networks (HPNs) are increasingly utilized to carry out such big data transfer. HPN management highly relies on an important capability of performance (mainly throughput) prediction to reserve sufficient bandwidth and meanwhile avoid over-provisioning that may result in unnecessary resource waste. This capability is critical to improving the resource (mainly bandwidth) utilization of dedicated connections and meeting various user requests for data transfer. Conventional methods conduct performance prediction by fitting prior observed transfer history with predefined loss functions, without considering unobservable latent factors such as competing loads on end hosts. Such latent factors also have a significant impact on the application-level data transfer performance, which may result in an inaccurate prediction model. In this paper, we first investigate the impact of latent factors and propose a clustering-based method to eliminate their negative impact on performance prediction. We then develop a robust machine learning-based performance predictor by: i) incorporating the proposed latent factor elimination method into data preprocessing, and ii) adopting a customized domain guided loss function. Extensive experimental results show that our predictor achieves significantly higher prediction accuracy than several other state-of-the-art methods.

Liu, Wuji↗

Performance Prediction of Big Data Transfer Through Experimental Analysis and Machine Learning

Big data transfer in next-generation scientific applications is now commonly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) through advance bandwidth reservation. To use HPN resources efficiently, provisioning agents need to carefully schedule data transfer requests and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized by the requesting user, could be simply wasted or cause extra overhead and complexity in management due to exclusive access. This calls for the capability of performance prediction to reserve bandwidth resources that match actual needs. Towards this goal, we employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements, which are collected over a span of several years from a large number of data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated HPN testbeds. We first identify a comprehensive list of attributes involved in a typical big data transfer process, including end host system configurations, network connection properties, and control parameters of data transfer methods. We then conduct an in-depth exploratory analysis of their impacts on application-level throughput, which provides insights into big data transfer performance and motivates the use of machine learning. We also investigate the applicability of machine learning algorithms and derive their general performance bounds for performance prediction of big data transfer in HPNs. Experimental results show that, with appropriate data preprocessing, the proposed machine learning-based approach achieves 95% or higher prediction accuracy in up to 90% of the cases with very noisy real-life performance measurements.

Yun, Daqing↗