Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45

Preliminary study of a serial-parallel redundant manipulator

The manipulator design discussed here results from the examination of some of the reasons why redundancy is necessary in general purpose manipulation systems. A spherical joint design actuated in-parallel, having the many advantages of parallel actuation, is described. In addition, the benefits of using redundant actuators are discussed and illustrated in the design by the elimination of loci of singularities from the usable workspace with the addition of only one actuator. Finally, what is known by the authors about space robotics requirements is summarized and the relevance of the proposed design matched against these requirements. The design problems outlined here are viewed as much from the mechanical engineering aspect as from concerns arising from the control and the programming of manipulators.

Hayward, Vincent↗

Parallel algorithms for computation of the manipulator inertia matrix

The development of an O(log2N) parallel algorithm for the manipulator inertia matrix is presented. It is based on the most efficient serial algorithm which uses the composite rigid body method. Recursive doubling is used to reformulate the linear recurrence equations which are required to compute the diagonal elements of the matrix. It results in O(log2N) levels of computation. Computation of the off-diagonal elements involves N linear recurrences of varying-size and a new method, which avoids redundant computation of position and orientation transforms for the manipulator, is developed. The O(log2N) algorithm is presented in both equation and graphic forms which clearly show the parallelism inherent in the algorithm.

Amin-Javaheri, Masoud↗

Hybrid interconnection structures for real-time parallel processing

The use of hybrid interconnection structures that combine link connections and bus connections for real-time parallel processing is discussed. Idealistic parallel computation models for two real-time computing applications are described with attention given to a tightly coupled network model for object tracking and a network model for image processing. Consideration is given to the following different interconnection structures: the crossbar, the hypercube, the circular linked array, and the bus array.

Kim, K. H.↗

ISEE 3 observations of solar wind thermal electrons with T-perpendicular greater than T-parallel

This study presents ISEE 3 observations of anomalous electron distributions for which T-perpendicular exceeds T-parallel in the solar wind near 1 AU. Twelve anomaly events were identified, lasting from 24 min to 6 hours. These events generally share the following characteristics: (1) high plasma density, (2) low solar wind speed, (3) magnetic field which is nearly transverse to the flow, and (4) low electron and ion temperatures. The processes of solar wind adiabatic expansion and isotropization via Coulomb collisions could be expected to lead to such anomalous anisotropies under conditions similar to those observed. However, these conditions actually produce T-perpendicular greater than T-parallel for only a small fraction of the time, suggesting that other mechanisms are also important in regulating solar wind electron distributions.

Phillips, J. L.↗

Parallelization of a three-dimensional compressible transition code

The compressible, three-dimensional, time-dependent Navier-Stokes equations are solved on a 20 processor Flex/32 computer. The code is a parallel implementation of an existing code operational on the Cray-2 at NASA Ames, which performs direct simulations of the initial stages of the transition process of wall-bounded flow at supersonic Mach numbers. Spectral collocation in all three spatial directions (Fourier along the plate and Chebyshev normal to it) ensures high accuracy of the flow variables. By hiding most of the parallelism in low-level routines, the casual user is shielded from most of the nonstandard coding constructs. Speedups of 13 out of a maximum of 16 are achieved on the largest computational grids.

Erlebacher, G.↗

Dynamic grid refinement for partial differential equations on parallel computers

The fast adaptive composite grid method (FAC) is an algorithm that uses various levels of uniform grids to provide adaptive resolution and fast solution of PDEs. An asynchronous version of FAC, called AFAC, that completely eliminates the bottleneck to parallelism is presented. This paper describes the advantage that this algorithm has in adaptive refinement for moving singularities on multiprocessor computers. This work is applicable to the parallel solution of two- and three-dimensional shock tracking problems.

Mccormick, S.↗

A parallel finite-difference method for computational aerodynamics

A finite-difference scheme for solving complex three-dimensional aerodynamic flow on parallel-processing supercomputers is presented. The method consists of a basic flow solver with multigrid convergence acceleration, embedded grid refinements, and a zonal equation scheme. Multitasking and vectorization have been incorporated into the algorithm. Results obtained include multiprocessed flow simulations from the Cray X-MP and Cray-2. Speedups as high as 3.3 for the two-dimensional case and 3.5 for segments of the three-dimensional case have been achieved on the Cray-2. The entire solver attained a factor of 2.7 improvement over its unitasked version on the Cray-2. The performance of the parallel algorithm on each machine is analyzed.

Swisshelm, Julie M.↗

Solving sparse triangular linear systems on parallel computers

This paper describes and compares three parallel algorithms for solving sparse triangular systems of equations. These methods involve some preprocessing overhead and are primarily of interest in solving many systems with the same coefficient matrix. The first approach is to use a fixed blocksize and form the inverse of the diagonal blocks. The second approach is to use a variable blocksize and reorder the unknowns so that the diagonal blocks are diagonal matrices. The latter technique is called level scheduling because of how it is represented in the adjacency graph, and both row-wise and jagged diagonal storage for the off-diagonal blocks are considered. These techniques are analyzed for general parallel computers and experiments are presented for the eight-processor Alliant FX/8.

Anderson, Edward↗

Magnetic pulsations at the quasi-parallel shock

The plasma and field properties of large-amplitude magnetic field pulsatins upstream from the quasi-parallel region of the earth's bow shock are examined in high time resolution using data from ISEE 1 and 2. The relative timing of the magnetic field profiles observed at the two spacecraft shows that some of the pulsations are convecting antisunward across the spacecraft while others are brief out/in motions of bow shock across the spacecraft. Pulsations with both timing signatures are the site of slowing and heating of the solar wind plasma. The ions tend to be only weakly heated in the convecting pulsations, while within the out/in pulsations the ion heating can be quite substantial but variable. This variation occurs not only from pulsation to pulsation but also from point to point within a given pulsation. In general, the hottest distributions within the out/in pulsations tend to occur in regions of lower density and field strength. Magnetic pulsations bear a number of similarities to previously identified hot diamagnetic cavity events as well as to more durable crossings of the quasi-parallel shock itself. These various phenomena may be different manifestations of the same basic physical processes, in particular the coupling of coherently reflected ions to the solar wind beam.

Thomsen, M. F.↗

Parallel flows with Soret effect in tilted cylinders

Henry and Roux (1986, 1987, 1988) have conducted extensive numerical studies on the interaction of Soret separation with convection in cylindrical geometry. Many of their solutions exhibit parallel flow away from end walls. Their parallel flow results can be matched by closed-form solutions. Solutions are nonunique in some parameter regions. Disappearance of one branch of solutions correlates with a sudden transition of Henry and Roux's results from a separated to a well-mixed flow.

Jacqmin, David↗

A parallel pipelined architecture for a digital multicarrier demodulator

A parallel pipelined architecture is presented for demultiplexing and demodulating SCPC/FDMA channels in real time. Specific algorithms are selected for each of the operations necessary for multicarrier demodulation. The selection is made based on their suitability for implementation into parallel-pipelined and sharing schemes. The demodulator is programmable and uses a single hardware module which is shared among all the channels for the recovery of clock, carrier, and data, resulting in large savings of power and hardware. The system is suitable for onboard processing of signals in satellites where power and area requirements are critical. The design is illustrated for the specific case of processing 800 FDMA channels at 64 kb/s each.

Fernandes, P. J.↗

A parallel row-based algorithm for standard cell placement with integrated error control

A new row-based parallel algorithm for standard-cell placement targeted for execution on a hypercube multiprocessor is presented. Key features of this implementation include a dynamic simulated-annealing schedule, row-partitioning of the VLSI chip image, and two novel approaches to control error in parallel cell-placement algorithms: (1) Heuristic Cell-Coloring; (2) Adaptive Sequence Length Control.

Sargent, Jeff S.↗

Analysis of a parallel multigrid algorithm

This paper considers the parallel multigrid algorithm of Frederickson and McBryan (1987). This algorithm uses multiple coarse-grid problems (instead of one) in the hope of accelerating convergence, and is found to have a close relationship to traditional multigrid methods. Specifically, the parallel coarse-grid correction operator is identical to a traditional multigrid coarse-grid correction operator, except that the mixing of high and low frequencies caused by aliasing error is removed. Appropriate relaxation operators can be chosen to take advantage of this property. Comparisons between the standard multigrid and the new method are made.

Chan, Tony F.↗

Production of electron conics by stochastic acceleration parallel to the magnetic field

Electron conics are enhancements in the electron flux at the edges of the electron loss cone. Such enhancements are a common feature in the electron distribution in the auroral zone. In analogy with ion conics, it has been suggested that electron conics are produced by waves which accelerate electrons perpendicular to the magnetic field. However, using a test particle simulation of the electron distribution it is shown that electron conics can be produced purely by stochastic acceleration of the electrons parallel to a dipole magnetic field. A possible wave mode that can produce parallel acceleration is the Alfven-ion cyclotron mode that has recently been shown to modulate the high energy part of the inverted-V electron distribution.

Temerin, Michael A.↗

Specularly reflected He(2+) at high Mach number quasi-parallel shocks

The first observations of near-specularly reflected He(2+) upstream from high Mach number quasi-parallel shocks are presented. The He(2+) observations are compared with simultaneous suprathermal proton observations to determine the density ratio of specularly reflected He(2+) to protons. The presence of specularly reflected He(2+) at high Mach number quasi-parallel shocks suggests that it may be a seed population for diffuse He(2+) seen in the same region.

Fuselier, S. A.↗

Merlin - Massively parallel heterogeneous computing

Hardware and software for Merlin, a new kind of massively parallel computing system, are described. Eight computers are linked as a 300-MIPS prototype to develop system software for a larger Merlin network with 16 to 64 nodes, totaling 600 to 3000 MIPS. These working prototypes help refine a mapped reflective memory technique that offers a new, very general way of linking many types of computer to form supercomputers. Processors share data selectively and rapidly on a word-by-word basis. Fast firmware virtual circuits are reconfigured to match topological needs of individual application programs. Merlin's low-latency memory-sharing interfaces solve many problems in the design of high-performance computing systems. The Merlin prototypes are intended to run parallel programs for scientific applications and to determine hardware and software needs for a future Teraflops Merlin network.

Wittie, Larry↗

The FORCE - A highly portable parallel programming language

This paper explains why the FORCE parallel programming language is easily portable among six different shared-memory multiprocessors, and how a two-level macro preprocessor makes it possible to hide low-level machine dependencies and to build machine-independent high-level constructs on top of them. These FORCE constructs make it possible to write portable parallel programs largely independent of the number of processes and the specific shared-memory multiprocessor executing them.

Jordan, Harry F.↗

Case studies in serial and parallel simulation

The design of a ground combat simulation in a time-driven and in an event-driven style is discussed. The differences between time-driven and event-driven simulation are illustrated, and performance results for the two systems executing both serially and in parallel are presented. It is noted that the event-driven style produces a more efficient system both serially and in parallel, primarily because it minimizes the number of messages produced and thereby reduces the amount of computation.

Wieland, Frederick↗