Data Movement Visualized - A Unified Performance Analysis Framework to Track Data Movement in Heterogeneous Architectures
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Pilot eye movement data by recording eye fixations on flight instruments
In this paper, we present a preliminary study of several different electronic data movement technologies. We detail our approach to classifying the technologies included in our study and present the preliminary results of some initial performance benchmarking. Our studies suggest that highly parallel TCP/IP streaming technologies, such as GridFTP and bbFTP, outperform commercial and open-source UDP-bursting technologies in several of the key data movement dimensions that we studied.
This milestone evaluates techniques to measure and, if possible, reduce data movement across all levels of the memory hierarchy, focusing on CPU/GPU page level data movement and on intra-GPU memory hierarchy. We quantitatively evaluate the efficacy of the techniques in reducing data movement and measure how performance tracks data movement reduction. We study a small collection of benchmarks and proxy mini-apps that run on advanced pre-exascale GPUs and on the Accelsim GPU simulator. Our approach has two thrusts: to measure advanced data movement reduction directives and techniques on the newest available GPUs, and to evaluate our benchmark set on simulated GPUs configured with architectural refinements to reduce data movement. We primarily evaluated NVidia-based architectures due to the unavailability of AMD GPU hardware and tools until very recently.
The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.
Eye movement data and other parameters including instrument readings, aircraft state and position variables, and control maneuvers were recorded while pilots flew ILS simulations in a B 737. The experiment itself employed seven airline pilots, each of whom flew approximately 40 approach/landing sequences. The simulator was equipped with a night visual scene but the scene was fogged out down to approximately 60 meters (200 ft). The instrument scanning appeared to follow aircraft parameters not physical position of instruments. One important implication of the results is: pilots look for categories or packets of information. Control inputs were tabulated according to throttle, wheel position, column, and pitch trim changes. Three seconds of eye movements before and after the control input were then obtained. Analysis of the eye movement data for the controlling periods showed clear patterns. The results suggest a set of miniscan patterns which are used according to the specific details of the situation. A model is developed which integrates scanning and controlling. Differentiations are made between monitoring and controlling scans.
The Large Hadron Collider (LHC) experiments distribute data by leveraging a diverse array of National Research and Education Networks (NRENs), where experiment data management systems treat networks as a “blackbox” resource. After the High Luminosity upgrade, the Compact Muon Solenoid (CMS) experiment alone will produce roughly 0.5 exabytes of data per year. NREN Networks are a critical part of the success of CMS and other LHC experiments. However, during data movement, NRENs are unaware of data priorities, importance, or need for quality of service, and this poses a challenge for operators to coordinate the movement of data and have predictable data flows across multi-domain networks. The overarching goal of SENSE (The Software-defined network for End-to-end Networked Science at Exascale) is to enable National Labs and universities to request and provision end-to-end intelligent network services for their application workflows leveraging SDN (Software-Defined Networking) capabilities. This work aims to allow LHC Experiments and Rucio, the data management software used by CMS Experiment, to allocate and prioritize certain data transfers over the wide area network. In this paper, we will present the current progress of the integration of SENSE, Multi-domain end-to-end SDN Orchestration with QoS (Quality of Service) capabilities, with Rucio, the data management software used by CMS Experiment.
The extreme-scale computing landscape is increasingly dominated by GPU-accelerated systems. At the same time, in-situ workflows that employ memory-to-memory inter-application data exchanges have emerged as an effective approach for leveraging these extreme-scale systems. In the case of GPUs, GPUDirect RDMA enables third-party devices, such as network interface cards, to access GPU memory directly and has been adopted for intra-application communications across GPUs. In this paper, we present an interoperable framework for GPU-based in-situ workflows that optimizes data movement using GPUDirect RDMA. Specifically, we analyze the characteristics of the possible data movement pathways between GPUs from an in-situ workflow perspective, and design a strategy that maximizes throughput. Furthermore, we implement this approach as an extension of the DataSpaces data staging service, and experimentally evaluate its performance and scalability on a current leadership GPU cluster. The performance results show that the proposed design reduces data-movement time by up to 53% and 40% for the sender and receiver, respectively, and maintains excellent scalability for up to 256 GPUs.
A problem often encountered when eye-movement measurement is conducted is the choice of 'indices' or 'statistics' available to present such information. The present study reports the use of the Time-Locked Time-History as a technique of value in the examination and presentation of eye-movement data. Plots created using this technique are labeled time-locked time-histories as they illustrate subject eye lookpoint during a period of time before and after a certain time-locking event. Events that occur with some degree of repetition, such as the onset or termination of control activities, warning signals, or changes in indicator positions may be utilized as time-locking events. The present study reports the use of this technique in an eye-movement study using a secondary task in which the subject must discriminate specific types of information in the display.
Neglecting the eccentric position of the eyes in the head can lead to erroneous interpretation of ocular motor data, particularly for near targets. We discuss the geometric effects that eye eccentricity has on the processing of target-directed eye and head movement data, and we highlight two approaches to processing and interpreting such data. The first approach involves determining the true position of the target with respect to the location of the eyes in space for evaluating the efficacy of gaze, and it allows calculation of retinal error directly from measured eye, head, and target data. The second approach effectively eliminates eye eccentricity effects by adjusting measured eye movement data to yield equivalent responses relative to a specified reference location (such as the center of head rotation). This latter technique can be used to standardize measured eye movement signals, enabling waveforms collected under different experimental conditions to be directly compared, both with the measured target signals and with each other. Mathematical relationships describing these approaches are presented for horizontal and vertical rotations, for both tangential and circumferential display screens, and efforts are made to describe the sensitivity of parameter variations on the calculated results.
Writing efficient parallel programs is complicated by the need to select the right data structure alignments and distributions, which determine the nature and volume of inter-processor communications. A large number of performance tools for parallel programs have been developed recently to expose these inter-processor communications. However, none of them support performance views or provide statistics in terms of inter-processor data structure interactions. A performance tool that tracks the interaction between individual data structures and the context of these interactions is essential for understanding the performance of both explicit message passing programs and data-parallel languages such as HPF. In this paper we discuss the use of compiler front end tools for automatically tracking data structure movements in message passing programs, and low-overhead monitoring and postprocessing of such codes. We demonstrate that robust instrumentation and low overhead monitoring of inter-processor data structure movements is possible, with the use of a number of NAS benchmark codes, run on the i860 hypercube. We also show that the data so collected can be used effectively by post processing tools that expose performance bottlenecks using graphical displays and performance statistics.
As supercomputers evolve, nodes are continually increasing in complexity. As a result, each generation of parallel systems brings new performance challenges. For instance, on recent systems inter-node communication has outperformed inter-socket, resulting in poor performance of many node-aware communication optimizations. Communication optimizations are critical for the performance and scalability of parallel applications, but are dependent on the parallel architecture, which varies significantly among recent generations of supercomputers. Furthermore, this paper investigates the performance of various paths of data movement on recent generations of systems, and analyzes the increased complexity of communication, particularly on recent heterogeneous systems. The paper also introduces MPI Advance, a communication library that enables optimizations to be created based on benchmark analysis of each emerging system.
Abstract Quantifying spatiotemporally explicit interactions within animal populations facilitates the understanding of social structure and its relationship with ecological processes. Data from animal tracking technologies (Global Positioning Systems [“GPS”]) can circumvent longstanding challenges in the estimation of spatiotemporally explicit interactions, but the discrete nature and coarse temporal resolution of data mean that ephemeral interactions that occur between consecutive GPS locations go undetected. Here, we developed a method to quantify individual and spatial patterns of interaction using continuous‐time movement models (CTMMs) fit to GPS tracking data. We first applied CTMMs to infer the full movement trajectories at an arbitrarily fine temporal scale before estimating interactions, thus allowing inference of interactions occurring between observed GPS locations. Our framework then infers indirect interactions—individuals occurring at the same location, but at different times—while allowing the identification of indirect interactions to vary with ecological context based on CTMM outputs. We assessed the performance of our new method using simulations and illustrated its implementation by deriving disease‐relevant interaction networks for two behaviorally differentiated species, wild pigs ( Sus scrofa ) that can host African Swine Fever and mule deer ( Odocoileus hemionus ) that can host chronic wasting disease. Simulations showed that interactions derived from observed GPS data can be substantially underestimated when temporal resolution of movement data exceeds 30‐min intervals. Empirical application suggested that underestimation occurred in both interaction rates and their spatial distributions. CTMM‐Interaction method, which can introduce uncertainties, recovered majority of true interactions. Our method leverages advances in movement ecology to quantify fine‐scale spatiotemporal interactions between individuals from lower temporal resolution GPS data. It can be leveraged to infer dynamic social networks, transmission potential in disease systems, consumer–resource interactions, information sharing, and beyond. The method also sets the stage for future predictive models linking observed spatiotemporal interaction patterns to environmental drivers.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
A processor includes a task scheduling unit and a compute unit coupled to the task scheduling unit. The task scheduling unit performs a task dependency assessment of a task dependency graph and task data requirements that correspond to each task of the plurality of tasks. Based on the task dependency assessment, the task scheduling unit schedules a first task of the plurality of tasks and a second proxy object of a plurality of proxy objects specified by the task data requirements such that a memory transfer of the second proxy object of the plurality of proxy objects occurs while the first task is being executed.
A processor includes a task scheduling unit and a compute unit coupled to the task scheduling unit. The task scheduling unit performs a task dependency assessment of a task dependency graph and task data requirements that correspond to each task of the plurality of tasks. Based on the task dependency assessment, the task scheduling unit schedules a first task of the plurality of tasks and a second proxy object of a plurality of proxy objects specified by the task data requirements such that a memory transfer of the second proxy object of the plurality of proxy objects occurs while the first task is being executed.