Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45

Solving sparse triangular linear systems on parallel computers

This paper describes and compares three parallel algorithms for solving sparse triangular systems of equations. These methods involve some preprocessing overhead and are primarily of interest in solving many systems with the same coefficient matrix. The first approach is to use a fixed blocksize and form the inverse of the diagonal blocks. The second approach is to use a variable blocksize and reorder the unknowns so that the diagonal blocks are diagonal matrices. The latter technique is called level scheduling because of how it is represented in the adjacency graph, and both row-wise and jagged diagonal storage for the off-diagonal blocks are considered. These techniques are analyzed for general parallel computers and experiments are presented for the eight-processor Alliant FX/8.

Anderson, Edward↗

Magnetic pulsations at the quasi-parallel shock

The plasma and field properties of large-amplitude magnetic field pulsatins upstream from the quasi-parallel region of the earth's bow shock are examined in high time resolution using data from ISEE 1 and 2. The relative timing of the magnetic field profiles observed at the two spacecraft shows that some of the pulsations are convecting antisunward across the spacecraft while others are brief out/in motions of bow shock across the spacecraft. Pulsations with both timing signatures are the site of slowing and heating of the solar wind plasma. The ions tend to be only weakly heated in the convecting pulsations, while within the out/in pulsations the ion heating can be quite substantial but variable. This variation occurs not only from pulsation to pulsation but also from point to point within a given pulsation. In general, the hottest distributions within the out/in pulsations tend to occur in regions of lower density and field strength. Magnetic pulsations bear a number of similarities to previously identified hot diamagnetic cavity events as well as to more durable crossings of the quasi-parallel shock itself. These various phenomena may be different manifestations of the same basic physical processes, in particular the coupling of coherently reflected ions to the solar wind beam.

Thomsen, M. F.↗

Parallel flows with Soret effect in tilted cylinders

Henry and Roux (1986, 1987, 1988) have conducted extensive numerical studies on the interaction of Soret separation with convection in cylindrical geometry. Many of their solutions exhibit parallel flow away from end walls. Their parallel flow results can be matched by closed-form solutions. Solutions are nonunique in some parameter regions. Disappearance of one branch of solutions correlates with a sudden transition of Henry and Roux's results from a separated to a well-mixed flow.

Jacqmin, David↗

A parallel pipelined architecture for a digital multicarrier demodulator

A parallel pipelined architecture is presented for demultiplexing and demodulating SCPC/FDMA channels in real time. Specific algorithms are selected for each of the operations necessary for multicarrier demodulation. The selection is made based on their suitability for implementation into parallel-pipelined and sharing schemes. The demodulator is programmable and uses a single hardware module which is shared among all the channels for the recovery of clock, carrier, and data, resulting in large savings of power and hardware. The system is suitable for onboard processing of signals in satellites where power and area requirements are critical. The design is illustrated for the specific case of processing 800 FDMA channels at 64 kb/s each.

Fernandes, P. J.↗

A parallel row-based algorithm for standard cell placement with integrated error control

A new row-based parallel algorithm for standard-cell placement targeted for execution on a hypercube multiprocessor is presented. Key features of this implementation include a dynamic simulated-annealing schedule, row-partitioning of the VLSI chip image, and two novel approaches to control error in parallel cell-placement algorithms: (1) Heuristic Cell-Coloring; (2) Adaptive Sequence Length Control.

Sargent, Jeff S.↗

Analysis of a parallel multigrid algorithm

This paper considers the parallel multigrid algorithm of Frederickson and McBryan (1987). This algorithm uses multiple coarse-grid problems (instead of one) in the hope of accelerating convergence, and is found to have a close relationship to traditional multigrid methods. Specifically, the parallel coarse-grid correction operator is identical to a traditional multigrid coarse-grid correction operator, except that the mixing of high and low frequencies caused by aliasing error is removed. Appropriate relaxation operators can be chosen to take advantage of this property. Comparisons between the standard multigrid and the new method are made.

Chan, Tony F.↗

Production of electron conics by stochastic acceleration parallel to the magnetic field

Electron conics are enhancements in the electron flux at the edges of the electron loss cone. Such enhancements are a common feature in the electron distribution in the auroral zone. In analogy with ion conics, it has been suggested that electron conics are produced by waves which accelerate electrons perpendicular to the magnetic field. However, using a test particle simulation of the electron distribution it is shown that electron conics can be produced purely by stochastic acceleration of the electrons parallel to a dipole magnetic field. A possible wave mode that can produce parallel acceleration is the Alfven-ion cyclotron mode that has recently been shown to modulate the high energy part of the inverted-V electron distribution.

Temerin, Michael A.↗

Specularly reflected He(2+) at high Mach number quasi-parallel shocks

The first observations of near-specularly reflected He(2+) upstream from high Mach number quasi-parallel shocks are presented. The He(2+) observations are compared with simultaneous suprathermal proton observations to determine the density ratio of specularly reflected He(2+) to protons. The presence of specularly reflected He(2+) at high Mach number quasi-parallel shocks suggests that it may be a seed population for diffuse He(2+) seen in the same region.

Fuselier, S. A.↗

Merlin - Massively parallel heterogeneous computing

Hardware and software for Merlin, a new kind of massively parallel computing system, are described. Eight computers are linked as a 300-MIPS prototype to develop system software for a larger Merlin network with 16 to 64 nodes, totaling 600 to 3000 MIPS. These working prototypes help refine a mapped reflective memory technique that offers a new, very general way of linking many types of computer to form supercomputers. Processors share data selectively and rapidly on a word-by-word basis. Fast firmware virtual circuits are reconfigured to match topological needs of individual application programs. Merlin's low-latency memory-sharing interfaces solve many problems in the design of high-performance computing systems. The Merlin prototypes are intended to run parallel programs for scientific applications and to determine hardware and software needs for a future Teraflops Merlin network.

Wittie, Larry↗

The FORCE - A highly portable parallel programming language

This paper explains why the FORCE parallel programming language is easily portable among six different shared-memory multiprocessors, and how a two-level macro preprocessor makes it possible to hide low-level machine dependencies and to build machine-independent high-level constructs on top of them. These FORCE constructs make it possible to write portable parallel programs largely independent of the number of processes and the specific shared-memory multiprocessor executing them.

Jordan, Harry F.↗

Case studies in serial and parallel simulation

The design of a ground combat simulation in a time-driven and in an event-driven style is discussed. The differences between time-driven and event-driven simulation are illustrated, and performance results for the two systems executing both serially and in parallel are presented. It is noted that the event-driven style produces a more efficient system both serially and in parallel, primarily because it minimizes the number of messages produced and thereby reduces the amount of computation.

Wieland, Frederick↗

Cross-fertilization between connectionist networks and highly parallel architectures

The theoretical and practical connections between connectionist schemes such as neural-network computers and traditional symbolic processing architectures involving a high degree of parallelism are explored, reviewing the results of recent investigations. Topics addressed include data flow, data structure, and control flow; conventional pointers; associative addressing; hashing and reduced representations; the problem of binding values to variables; and levels of parallelism. It is concluded that connectionism is more closely related to traditional computer science and technology than is generally admitted; more cooperation between followers of the two approaches is recommended.

Barnden, John↗

Adaptive domain decomposition for Monte Carlo simulations on parallel processors

A method is described for performing direct simulation Monte Carlo (DSMC) calculations on parallel processors using adaptive domain decomposition to distribute the computational work load. The method has been implemented on a commercially available hypercube and benchmark results are presented which show the performance of the method relative to current supercomputers. The problems studied were simulations of equilibrium conditions in a closed, stationary box, a two-dimensional vortex flow, and the hypersonic, rarefield flow in a two-dimensional channel. For these problems, the parallel DSMC method ran 5 to 13 times faster than on a single processor of a Cray-2. The adaptive decomposition method worked well in uniformly distributing the computational work over an arbitrary number of processors and reduced the average computational time by over a factor of two in certain cases.

Wilmoth, Richard G.↗

An empirical study of FORTRAN programs for parallelizing compilers

Some results are reported from an empirical study of program characteristics that are important in parallelizing compiler writers, especially in the area of data dependence analysis and program transformations. The state of the art in data dependence analysis and some parallel execution techniques are examined. The major findings are included. Many subscripts contain symbolic terms with unknown values. A few methods of determining their values at compile time are evaluated. Array references with coupled subscripts appear quite frequently; these subscripts must be handled simultaneously in a dependence test, rather than being handled separately as in current test algorithms. Nonzero coefficients of loop indexes in most subscripts are found to be simple: they are either 1 or -1. This allows an exact real-valued test to be as accurate as an exact integer-valued test for one-dimensional or two-dimensional arrays. Dependencies with uncertain distance are found to be rather common, and one of the main reasons is the frequent appearance of symbolic terms with unknown values.

Shen, Zhiyu↗

Optimal parallel evaluation of AND trees

A quantitative analysis based on both preemptive and nonpreemptive critical-path scheduling algorithms is presently conducted for the optimal degree of parallelism required in evaluating a given AND tree. The optimal degree of parallelism is found to depend on problem complexity, precedence-graph shape, and task-time distribution along each path. In addition to demonstrating the optimality of the preemptive critical-path scheduling algorithm for evaluating an arbitrary AND tree on a fixed number of processors, the possibility of efficiently ascertaining tight bounds on the number of processors for optimal processor-time efficiency is illustrated.

Wah, Benjamin W.↗

Fast Parallel Computation Of Manipulator Inverse Dynamics

Method for fast parallel computation of inverse dynamics problem, essential for real-time dynamic control and simulation of robot manipulators, undergoing development. Enables exploitation of high degree of parallelism and, achievement of significant computational efficiency, while minimizing various communication and synchronization overheads as well as complexity of required computer architecture. Universal real-time robotic controller and simulator (URRCS) consists of internal host processor and several SIMD processors with ring topology. Architecture modular and expandable: more SIMD processors added to match size of problem. Operate asynchronously and in MIMD fashion.

Fijany, Amir↗

Dynamic Analysis and Control of Lightweight Manipulators with Flexible Parallel Link Mechanisms

The objective is the theoretical analysis and the experimental verification of dynamics and control of a two link flexible manipulator with a flexible parallel link mechanism. Nonlinear equations of motion of the lightweight manipulator are derived by the Lagrangian method in symbolic form to better understand the structure of the dynamic model. The resulting equation of motion have a structure which is useful to reduce the number of terms calculated, to check correctness, or to extend the model to higher order. A manipulator with a flexible parallel link mechanism is a constrained dynamic system whose equations are sensitive to numerical integration error. This constrained system is solved using singular value decomposition of the constraint Jacobian matrix. Elastic motion is expressed by the assumed mode method. Mode shape functions of each link are chosen using the load interfaced component mode synthesis. The discrepancies between the analytical model and the experiment are explained using a simplified and a detailed finite element model.

Lee, Jeh Won↗

Inflated speedups in parallel simulations via malloc()

Discrete-event simulation programs make heavy use of dynamic memory allocation in order to support simulation's very dynamic space requirements. When programming in C one is likely to use the malloc() routine. However, a parallel simulation which uses the standard Unix System V malloc() implementation may achieve an overly optimistic speedup, possibly superlinear. An alternate implementation provided on some (but not all systems) can avoid the speedup anomaly, but at the price of significantly reduced available free space. This is especially severe on most parallel architectures, which tend not to support virtual memory. It is shown how a simply implemented user-constructed interface to malloc() can both avoid artificially inflated speedups, and make efficient use of the dynamic memory space. The interface simply catches blocks on the basis of their size. The problem is demonstrated empirically, and the effectiveness of the solution is shown both empirically and analytically.

Nicol, David M.↗