Search NASA⌕ Search

SEARCH · Search NASA

Results for “HADOOP”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Rendezvous algorithms for large-scale modeling and simulation

Rendezvous algorithms encode a communication pattern that is useful when processors sending data do not know who the receiving processors should be, or vice versa. The idea is to define an intermediate decomposition where datums from different sending processors can ”rendezvous” to perform a computation, in a manner that both the senders and eventual receivers of the results can identify the appropriate rendezvous processor. Though they were originally designed for interpolating between overlaid grids with independent parallel decompositions (Plimpton et al., 2004), we have recently found rendezvous algorithms useful for a variety of operations in particle- or grid-based simulation codes when running large problems on large numbers of processors. In particular, we show they can perform well when a load-balanced intermediate decomposition is randomized and not spatial, requiring all-to-all communication to move data between processors. In this case rendezvous algorithms leverage the large bisection communication bandwidths which parallel machines provide. We describe how rendezvous algorithms work in a scientific computing context and give specific examples for molecular dynamics and Direct Simulation Monte Carlo codes which result in dramatic performance improvements versus simpler algorithms which do not scale as well. We explain how a generic rendezvous algorithm can be implemented, and also point out similarities with the MapReduce paradigm popularized by Google and Hadoop.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

Online Power System Event Detection via Bidirectional Generative Adversarial Networks

Accurate and speedy detection of power system events is critical to enhancing the reliability and resiliency of power systems. Although supervised deep learning algorithms show great promise in identifying power system events, they require a large volume of high-quality event labels for training. This paper develops a bidirectional anomaly generative adversarial network (GAN)-based algorithm to detect power system events using streaming PMU data, which does not rely on a huge amount of event labels. By introducing conditional entropy constraint in the objective function of GAN and graph signal processing-based PMU sorting technique, our proposed algorithm significantly outperforms state-of-the-art event detection algorithms in terms of accuracy. To facilitate the adoption of the proposed algorithm, a prototype online platform is also developed using Apache Hadoop, Kafka, and Spark to enable real-time event detection. Here, the accuracy and computational efficiency of the proposed algorithm are validated using a large-scale real-world PMU dataset from the Eastern Interconnection of the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Offloading Calculations to Computational Storage Devices: Spark and HDFS [Slides]

The objective is to evaluate the capabilities of multiple CSDs (provided by NDG Systems) using Hadoop Filesystem and Apache Spark. The independent variables are: number of CSDs, 0, 1, 2, 4, or 6; size of dataset, 1 GB, 5 GB, 10 GB; type of dataset, one large file with all of the data, 10 files, 100 files. The dependent variables are: job time; execution time. the constants are operations on the dataset.

97 MATHEMATICS AND COMPUTING↗