Search NASA⌕ Search

SEARCH · Search NASA

Results for “data dependencies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis

Checkpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 13 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint.

HPC↗

An efficient data dependence analysis for parallelizing compilers

A novel algorithm, called the lambda test, is presented for an efficient and accurate data dependence analysis of multidimensional array references. It extends the numerical methods to allow all dimensions of array references to be tested simultaneously. Hence, it combines the efficiency and the accuracy of the both approaches. This algorithm has been implemented in PARAFRASE, a FORTRAN program parallelization restructurer developed at the University of Illinois at Urbana-Champaign. Some experimental results are presented to show its effectiveness.

Li, Zhiyuan↗

Data dependent systems methodology for lumped mass modeling of structures

Limitations of the frequency domain methods in analyzing structura1 vibrations has created an awareness of the comparative merits of the time domain methods. Although time domain methods would be ideal for modeling large precisions space systems, the popular methods based on fitting theoretical response to actual data by least squares are too sensitive to noise and require too much data to be suitable for orbiting space crafts. This paper briefly reviews the theory and illustrative applications of a time domain methodology called Data Dependent Systems (DDS) that eliminates these limitations. Simulation results are presented to demonstrate a better than 4-place accuracy in the identifications of all system parameters, both modal (frequencies, damping ratios, and mode shapes) and physical (mass, stiffness, and damping matrices).

Pandit, Sudhakar M.↗

Minimizing inner product data dependencies in conjugate gradient iteration

The amount of concurrency available in conjugate gradient iteration is limited by the summations required in the inner product computations. The inner product of two vectors of length N requires time c log(N), if N or more processors are available. This paper describes an algebraic restructuring of the conjugate gradient algorithm which minimizes data dependencies due to inner product calculations. After an initial start up, the new algorithm can perform a conjugate gradient iteration in time c*log(log(N)).

Vanrosendale, J.↗

Trends in column ozone based on TOMS data - Dependence on month, latitude, and longitude

On the basis of the TOMS satellite column ozone data in latitudes 70 deg S-70 deg N from November 1978 to May 1990, a statistical model is used to estimate the trends in ozone as a function of latitude, longitude, and month. The trends in the TOMS ozone data are highly seasonal and dependent on location. Near the equator, the estimated monthly trends are not significantly different from zero. For high latitudes, most of the estimated monthly trends are negative. In January, February, and March, there are some positive trend estimates in the western hemisphere around latitude 60 deg N. The most negative trends for these three months also appear in the high latitudes of the northern hemisphere. Starting in June, more negative trends appear in the latitudes 50 deg S-70 deg S than the trends in the rest of the world considered. A large depletion develops during the spring time (September to November) in the southern high-latitude region, and the area of peak ozone decline is moving eastward during the period. The largest negative trends (about -29 percent per decade) for the area considered in this study appear in October around the latitude 70 deg S and longitudes 20 deg W-100 deg W region. For the northern hemisphere, the year-round trend estimates for latitudes 30 deg N-70 deg N range from -0.96 percent to -7.43 percent per decade. In the latitudes 30 deg N-50 deg N, the winter trend estimates are more negative than those for the summer and the fall. However, this pattern did not hold for latitudes 50 deg N-70 deg N.

Niu, Xufeng↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

Strategies for weather-dependent data acquisition

A strategy for data acquisition from a very distant spacecraft is presented, when the system performance can be severely degraded by the Earth's weather due to the high microwave frequency being used. When there is no minimum rate to be maintained, the optimum strategy is the greedy strategy, which always transmits at the single rate which maximizes the expected data returned. If there is a minimum data rate, the optimum strategy transmits simultaneously at the minimum or base data rate and at a bonus data rate. A coding system designed for the bandwidth-constrained degraded broadcast channels used. The optimum version of this system can, under realistic assumptions, save on the order of 5 dB over the conservative strategy of just transmitting at a single lower data rate.

Posner, E. C.↗

Strategies for weather-dependent data acquisition

A strategy for data acquisition from a very distant spacecraft is presented, when the system performance can be severely degraded by the Earth's weather due to the high microwave frequency being used. When there is no minimum rate to be maintained, the optimum strategy is the greedy strategy, which always transmits at the single rate which maximizes the expected data returned. If there is a minimum data rate, tne optimum strategy transmits simultaneously at the minimum or base data rate and at a bonus data rate. A coding system designed for the bandwidth-constrained degraded broadcast channels used. The optimum version of this system can, under realistic assumptions, save on the order of 5 dB over the conservative strategy of just transmitting at a single lower data rate. Previously announced in STAR as N82-11286

Posner, E. C.↗

An expert system to analyze high frequency dependent data for the space shuttle main engine turbopumps

The prototype expert system ADDAMX identifies selected sinusoid frequencies from spectral data graphs as speed frequencies and harmonics from each turbopump, frequency feed through from one turbopump to another, frequencies generated by turbopump bearings, pseudo 3N for the phase 2 high pressure fuel turbopump, and electrical noise. ADDAMX does the analysis in an interactive or batch mode and the results can be displayed on the screen or hardcopy.

Garcia, Raul C., Jr.↗

POSTER: Automatic Differentiation of Parallel Loops with Formal Methods

The accompanying poster to this short paper presents a combination of reverse mode AD and formal methods to enable efficient differentiation of (or backpropagation through) shared-memory parallel code. Compared to the state of the art, our approach can more often avoid the need for atomic updates or private data copies during the parallel derivative computation, even in the presence of unstructured or data-dependent data access patterns. This is achieved by gathering information about the memory access patterns from the input program, which is assumed to be correctly parallelized. This information is then used to build a model of assertions in a theorem prover, which can be used to check the safety of shared memory accesses during the parallel derivative computation

Automatic Differentiation↗

Revised radiometric calibration technique for LANDSAT-4 Thematic Mapper data

Depending on detector number, there are random fluctuations in the background level for spectral band 1 of magnitudes ranging from 2 to 3.5 digital numbers (DN). Similar variability is observed in all the other reflective bands, but with smaller magnitude in the range 0.5 to 2.5 DN. Observations of background reference levels show that line dependent variations in raw TM image data and in the associated calibration data can be measured and corrected within an operational environment by applying simple offset corrections on a line-by-line basis. The radiometric calibration procedure defined by the Canadian Center for Remote Sensing was revised accordingly in order to prevent striping in the output product.

Murphy, J.↗

Automatic Data Distribution for CFD Applications on Structured Grids

Development of HPF versions of NPB and ARC3D showed that HPF has potential to be a high level language for parallelization of CFD applications. The use of HPF requires an intimate knowledge of the applications and a detailed analysis of data affinity, data movement and data granularity. Since HPF hides data movement from the user even with this knowledge it is easy to overlook pieces of the code causing low performance of the application. In order to simplify and accelerate the task of developing HPF versions of existing CFD applications we have designed and partially implemented ADAPT (Automatic Data Distribution and Placement Tool). The ADAPT analyzes a CFD application working on a single structured grid and generates HPF TEMPLATE, (RE)DISTRIBUTION, ALIGNMENT and INDEPENDENT directives. The directives can be generated on the nest level, subroutine level, application level or inter application level. ADAPT is designed to annotate existing CFD FORTRAN application performing computations on single or multiple grids. On each grid the application can considered as a sequence of operators each applied to a set of variables defined in a particular grid domain. The operators can be classified as implicit, having data dependences, and explicit, without data dependences. In order to parallelize an explicit operator it is sufficient to create a template for the domain of the operator, align arrays used in the operator with the template, distribute the template, and declare the loops over the distributed dimensions as INDEPENDENT. In order to parallelize an implicit operator, the distribution of the operator's domain should be consistent with the operator's dependences. Any dependence between sections distributed on different processors would preclude parallelization if compiler does not have an ability to pipeline computations. If a data distribution is "orthogonal" to the dependences of an implicit operator then the loop which implements the operator can be declared as INDEPENDENT.

Frumkin, Michael↗

Parafrase restructuring of FORTRAN code for parallel processing

Parafrase transforms a FORTRAN code, subroutine by subroutine, into a parallel code for a vector and/or shared-memory multiprocessor system. Parafrase is not a compiler; it transforms a code and provides information for a vector or concurrent process. Parafrase uses a data dependency to reveal parallelism among instructions. The data dependency test distinguishes between recurrences and statements that can be directly vectorized or parallelized. A number of transformations are required to build a data dependency graph.

Wadhwa, Atul↗

DynPaC: Coarse-Grained, Dynamic, and Partially Reconfigurable Array for Streaming Applications

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across kernels. However, streaming applications often are data-dependent, leading to variable kernel execution times depending on the input data and impacting the throughput of the entire pipeline if resources are statically allocated. Therefore, in this paper, we discuss the design of DynPaC — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We discuss the required software and hardware components to manage partial dynamic reconfiguration. We demonstrate that by supporting partial dynamic reconfiguration, we can obtain an average speedup of 1.44X for a representative set of applications w.r.t. static partitioning, with a limited area overhead (6.4% of the entire chip).

Tan, Cheng↗

A Tool for Automatic Data Distribution for CFD Applications on Structured Grids

Development of HPF versions of NPB and ARC3D has shown that HPF provides an efficient, concise way to express parallelism and to organize data traffic. The use of HPF, as noted in the papers, requires an intimate knowledge of the applications and a detailed analysis of data affinity, data movement, and data granularity. To simplify and accelerate the task of developing HPF versions of existing CFD applications we have designed and implemented ADAPT (Automatic Data Alignment and Placement Tool). ADAPT analyzes a CFD application working on a single structured grid and generates HPF TEMPLATE, (RE)DISTRIBUTION, ALIGNMENT, and INDEPENDENT directives. The directives can be generated on the nest level, subroutine level, application level, or on the application interface level. ADAPT annotates an existing CFD FORTRAN application, performing computations on single or multiple grids. On each grid the application is considered as a sequence of operators, each applied to a set of variables defined in a particular grid domain. ADAPT automatically detects implicit operators (i.e., having data dependences) and explicit operators (without data dependences). For parallelization of an explicit operator ADAPT creates a template for the operator domain, aligns arrays used in the operator with the template, distributes the template, and declares the loops over the distributed dimensions as INDEPENDENT. For parallelization of an implicit operator, the distribution of the operator's domain should be consistent with the operator's dependences. Any dependence between sections distributed on different processors would preclude parallelization if the compiler does not have an ability to pipeline computations. If a data distribution is "orthogonal" to the dependences of an implicit operator, then the loop which implements the operator can be declared as INDEPENDENT. ADAPT starts with an analysis of array index expressions of the loop nests. For each pair of arrays referenced in an assignment statement, it generates an arc in the alignment graph and annotates it with an affinity relation. The template, alignment, and distribution directives for a particular loop nest are then derived from a transitive closure of the affinity relation. A compromise of data distributions in different nests and subroutines is achieved by merging annotated alignment graphs for adjacent nests/stibroutine calls in the nest/call graph of the application in the process called distribution lifting. ADAPT has been implemented as a C++ program running in conjunction with a parallelization tool called CAPTools. ADAPT uses the parse tree, interprocedural analysis and application database generated by CAPTools. It also uses the Directed Graph class, initially implemented in p2d2 (parallel debugger oi distributed programs), and some other classes supporting symbolic computations. ADAPT uses data distribution techniques described. ADAPT was tested with ARC3D and the FT benchmark and has demonstrated a code performance within a factor of 1.5 of handwritten versions.

Frumkin, Michael↗