Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61

Massively parallel processor

A brief description is given of the Massively Parallel Processor (MPP). Major applications of the MPP are in the area of image processing (where the operands are often very small integers) from very high spatial resolution passive image sensors, signal processing of radar data, and numerical modeling simulations of climate. The system can be programmed in assembly language or a high level language. Information on background, status, architecture, programming, hardware reliability, applications, and the MPP's development as a national resource for parallel algorithm research are presented in outline form.

Source record↗

Application of powder densification models to the consolidation processing of composites

Unidirectional fiber reinforced metal matrix composite tapes (containing a single layer of parallel fibers) can now be produced by plasma deposition. These tapes can be stacked and subjected to a thermomechanical treatment that results in a fully dense near net shape component. The mechanisms by which this consolidation step occurs are explored, and models to predict the effect of different thermomechanical conditions (during consolidation) upon the kinetics of densification are developed. The approach is based upon a methodology developed by Ashby and others for the simpler problem of HIP of spherical powders. The complex problem is devided into six, much simpler, subproblems, and then their predicted contributions are added to densification. The initial problem decomposition is to treat the two extreme geometries encountered (contact deformation occurring between foils and shrinkage of isolated, internal pores). Deformation of these two geometries is modelled for plastic, power law creep and diffusional flow. The results are reported in the form of a densification map.

Wadley, H. N. G.↗

Optoelectronic Tool Adds Scale Marks to Photographic Images

A simple, easy-to-use optoelectronic tool projects scale marks that become incorporated into photographic images (including film and electronic images). The sizes of objects depicted in the images can readily be measured by reference to the scale marks. The role played by the scale marks projected by this tool is the same as that of the scale marks on a ruler placed in a scene for the purpose of establishing a length scale. However, this tool offers the advantage that it can put scale marks quickly and safely in any visible location, including a location in which placement of a ruler would be difficult, unsafe, or time-consuming. The tool (see Figure 1) includes an aluminum housing, within which are mounted four laser diodes that operate at a wavelength of 670 nm. The laser diodes are spaced 1 in. (2.54 cm) apart along a baseline. The laser diodes are mounted with setscrews, which are used to adjust their beams to make them all parallel to each other and perpendicular to the baseline. During the adjustment process, the effect of the adjustments is observed by measuring the positions of the laser-beam spots on a target 80 ft (approx.24 m) away. Once the adjustments have been completed, the laser beams define three 1-in. (2.54-cm) intervals and the location of each beam is defined to within 1/16 in. (approx.1.6 mm) at any target distance out to about 80 ft (approx.24 m). The distance between the laser-beam spots as seen in an image is strictly defined only along an axis parallel to the baseline and perpendicular to the laser beam (also perpendicular to the line of sight of the camera, assuming that the camera-to-target distance is much greater than the distance between the tool and the camera lens). If a flat target surface illuminated by the laser beams is tilted with respect to the aforesaid axis, then the distance along the target surface between scale marks is proportional to the secant of the tilt angle. If one knows the tilt angle, one can correct for it. Even if one does not know the tilt angle precisely, it may not matter: For example, at a tilt of 10 , the secant is approximately 1.0154, so that the tilt error is only about 1.54 percent, which is negligibly small for a typical application in which only approximate measurements are needed.

Stevenson, Charlie↗

Vanishing dual-task interference after practice: has the bottleneck been eliminated or is it merely latent?

Practice can, in some cases, largely eliminate measured dual-task interference. Does this absence of interference indicate the absence of a processing bottleneck (defined as an inability to carry out certain stages in parallel)? The authors show that a bottleneck need not produce any observable interference, provided that there is no temporal overlap in the demand for bottleneck stages on the 2 tasks. Such a "latent" bottleneck is especially likely after practice, when central stages are short. The authors provide new evidence that a latent bottleneck occurred for a participant who produced no interference in M. Van Selst, E. Ruthruff, and J. C. Johnston (1999). These findings demonstrate that the absence of dual-task interference does not necessarily indicate the absence of a processing bottleneck.

Clinical Trial↗

Performance Assessment of OVERFLOW on Distributed Computing Environment

The aerodynamic computer code, OVERFLOW, with a multi-zone overset grid feature, has been parallelized to enhance its performance on distributed and shared memory paradigms. Practical application benchmarks have been set to assess the efficiency of code's parallelism on high-performance architectures. The code's performance has also been experimented with in the context of the distributed computing paradigm on distant computer resources using the Information Power Grid (IPG) toolkit, Globus. Two parallel versions of the code, namely OVERFLOW-MPI and -MLP, have developed around the natural coarse grained parallelism inherent in a multi-zonal domain decomposition paradigm. The algorithm invokes a strategy that forms a number of groups, each consisting of a zone, a cluster of zones and/or a partition of a large zone. Each group can be thought of as a process with one or multithreads assigned to it and that all groups run in parallel. The -MPI version of the code uses explicit message-passing based on the standard MPI library for sending and receiving interzonal boundary data across processors. The -MLP version employs no message-passing paradigm; the boundary data is transferred through the shared memory. The -MPI code is suited for both distributed and shared memory architectures, while the -MLP code can only be used on shared memory platforms. The IPG applications are implemented by the -MPI code using the Globus toolkit. While a computational task is distributed across multiple computer resources, the parallelism can be explored on each resource alone. Performance studies are achieved with some practical aerodynamic problems with complex geometries, consisting of 2.5 up to 33 million grid points and a large number of zonal blocks. The computations were executed primarily on SGI Origin 2000 multiprocessors and on the Cray T3E. OVERFLOW's IPG applications are carried out on NASA homogeneous metacomputing machines located at three sites, Ames, Langley and Glenn. Plans for the future will exploit the distributed parallel computing capability on various homogeneous and heterogeneous resources and large scale benchmarks. Alternative IPG toolkits will be used along with sophisticated zonal grouping strategies to minimize the communication time across the computer resources.

Djomehri, M. Jahed↗

I/O-Efficient Scientific Computation Using TPIE

In recent years, input/output (I/O)-efficient algorithms for a wide variety of problems have appeared in the literature. However, systems specifically designed to assist programmers in implementing such algorithms have remained scarce. TPIE is a system designed to support I/O-efficient paradigms for problems from a variety of domains, including computational geometry, graph algorithms, and scientific computation. The TPIE interface frees programmers from having to deal not only with explicit read and write calls, but also the complex memory management that must be performed for I/O-efficient computation. In this paper we discuss applications of TPIE to problems in scientific computation. We discuss algorithmic issues underlying the design and implementation of the relevant components of TPIE and present performance results of programs written to solve a series of benchmark problems using our current TPIE prototype. Some of the benchmarks we present are based on the NAS parallel benchmarks while others are of our own creation. We demonstrate that the central processing unit (CPU) overhead required to manage I/O is small and that even with just a single disk, the I/O overhead of I/O-efficient computation ranges from negligible to the same order of magnitude as CPU time. We conjecture that if we use a number of disks in parallel this overhead can be all but eliminated.

Vengroff, Darren Erik↗

A Multi-Level Parallelization Concept for High-Fidelity Multi-Block Solvers

The integration of high-fidelity Computational Fluid Dynamics (CFD) analysis tools with the industrial design process benefits greatly from the robust implementations that are transportable across a wide range of computer architectures. In the present work, a hybrid domain-decomposition and parallelization concept was developed and implemented into the widely-used NASA multi-block Computational Fluid Dynamics (CFD) packages implemented in ENSAERO and OVERFLOW. The new parallel solver concept, PENS (Parallel Euler Navier-Stokes Solver), employs both fine and coarse granularity in data partitioning as well as data coalescing to obtain the desired load-balance characteristics on the available computer platforms. This multi-level parallelism implementation itself introduces no changes to the numerical results, hence the original fidelity of the packages are identically preserved. The present implementation uses the Message Passing Interface (MPI) library for interprocessor message passing and memory accessing. By choosing an appropriate combination of the available partitioning and coalescing capabilities only during the execution stage, the PENS solver becomes adaptable to different computer architectures from shared-memory to distributed-memory platforms with varying degrees of parallelism. The PENS implementation on the IBM SP2 distributed memory environment at the NASA Ames Research Center obtains 85 percent scalable parallel performance using fine-grain partitioning of single-block CFD domains using up to 128 wide computational nodes. Multi-block CFD simulations of complete aircraft simulations achieve 75 percent perfect load-balanced executions using data coalescing and the two levels of parallelism. SGI PowerChallenge, SGI Origin 2000, and a cluster of workstations are the other platforms where the robustness of the implementation is tested. The performance behavior on the other computer platforms with a variety of realistic problems will be included as this on-going study progresses.

Hatay, Ferhat F.↗

Extended-Range Ultrarefractive 1D Photonic Crystal Prisms

A proposal has been made to exploit the special wavelength-dispersive characteristics of devices of the type described in One-Dimensional Photonic Crystal Superprisms (NPO-30232) NASA Tech Briefs, Vol. 29, No. 4 (April 2005), page 10a. A photonic crystal is an optical component that has a periodic structure comprising two dielectric materials with high dielectric contrast (e.g., a semiconductor and air), with geometrical feature sizes comparable to or smaller than light wavelengths of interest. Experimental superprisms have been realized as photonic crystals having three-dimensional (3D) structures comprising regions of amorphous Si alternating with regions of SiO2, fabricated in a complex process that included sputtering. A photonic crystal of the type to be exploited according to the present proposal is said to be one-dimensional (1D) because its contrasting dielectric materials would be stacked in parallel planar layers; in other words, there would be spatial periodicity in one dimension only. The processes of designing and fabricating 1D photonic crystal superprisms would be simpler and, hence, would cost less than do those for 3D photonic crystal superprisms. As in 3D structures, 1D photonic crystals may be used in applications such as wavelength-division multiplexing. In the extended-range configuration, it is also suitable for spectrometry applications. As an engineered structure or artificially engineered material, a photonic crystal can exhibit optical properties not commonly found in natural substances. Prior research had revealed several classes of photonic crystal structures for which the propagation of electromagnetic radiation is forbidden in certain frequency ranges, denoted photonic bandgaps. It had also been found that in narrow frequency bands just outside the photonic bandgaps, the angular wavelength dispersion of electromagnetic waves propagating in photonic crystal superprisms is much stronger than is the angular wavelength dispersion obtained by use of conventional prisms and diffraction gratings and is highly nonlinear.

Ting, David Z.↗

Improved silicon carbide for advanced heat engines

The development of silicon carbide materials of high strength was initiated and components of complex shape and high reliability were formed. The approach was to adapt a beta-SiC powder and binder system to the injection molding process and to develop procedures and process parameters capable of providing a sintered silicon carbide material with improved properties. The initial effort was to characterize the baseline precursor materials, develop mixing and injection molding procedures for fabricating test bars, and characterize the properties of the sintered materials. Parallel studies of various mixing, dewaxing, and sintering procedures were performed in order to distinguish process routes for improving material properties. A total of 276 modulus-of-rupture (MOR) bars of the baseline material was molded, and 122 bars were fully processed to a sinter density of approximately 95 percent. Fluid mixing techniques were developed which significantly reduced flaw size and improved the strength of the material. Initial MOR tests indicated that strength of the fluid-mixed material exceeds the baseline property by more than 33 percent. the baseline property by more than 33 percent.

Whalen, Thomas J.↗

Magnetopause and boundary layer

A brief overview is given of our present knowledge, observational and theoretical, of the structure of the magnetopause and the adjoining plasma boundary layer. Particular attention is given to the relationship between these electromagnetic and plasma structures on the front lobe of the magnetosphere and the magnetic field reconnection process. Items discussed include: magnetopause thickness; behavior of magnetic field components parallel and perpendicular to the magnetopause; particle energization; structure of the boundary layer from reconnection theory.

Sonnerup, B. U. O.↗

Beyond the supercomputer

A NASA-directed development of massively parallel processor (MPP) computers is outlined, noting intended applications for data processing for near term earth resource and environment mapping, radar, and television transmissions. The MPP is designed to perform 100 billion operations/sec to obtain satisfactory image processing, while separate processing units correct distortions, register images, calculate correlation functions, and classify multispectral characteristics. Arrays of 1s and 0s will be manipulated in analog-to-digital conversions generating separate planes corresponding to powers of binaries. Data wires are replaced by fiber-optic tubes or thousands of wires, and single logic gates are replaced by thousands of logic gates and every memory element by thousands of memory elements. Features of the interconnections and the images control processor units are detailed, along with implementation of sliders for program flexibility.

Schaefer, D. H.↗

Real-time data compressor for Eos-class missions

A conceptual design for a real-time VLSI compressor capable of processing rate up to one gigabit per second is presented. This scheme is capable of providing a three-to-one distortion-free data reduction factor to both the High Resolution Imaging Spectrometer and processed SAR imaging data. The design uses a VLSI parallel/piplined architecture capable of processing at a real time rate. The design consists of a parallel array of VLSI compressor modules. Each module is built on a single customized VLSI chip using existing state-of-the-art semiconductor technology.

Lee, Jun-Ji↗

Excitation of whistlers and waves with mixed polarization by newborn cometary ions

The present study has been motivated by the ICE wave measurements. It is found that the newborn cometary ions, particularly the protons, can excite whistlers and waves with frequencies much higher than the proton gyrofrequency but with mixed electrostatic and electromagnetic polarization. For the case of oblique propagation the newborn ions are treated as if they are unmagnetized. This is justified not only because the wave frequencies are high but also because the growth rates are large. On the other hand in the case of parallel propagation the growth rate is much smaller, and the excitation process seems to be unimportant.

Wu, C. S.↗

A tesselated probabilistic representation for spatial robot perception and navigation

The ability to recover robust spatial descriptions from sensory information and to efficiently utilize these descriptions in appropriate planning and problem-solving activities are crucial requirements for the development of more powerful robotic systems. Traditional approaches to sensor interpretation, with their emphasis on geometric models, are of limited use for autonomous mobile robots operating in and exploring unknown and unstructured environments. Here, researchers present a new approach to robot perception that addresses such scenarios using a probabilistic tesselated representation of spatial information called the Occupancy Grid. The Occupancy Grid is a multi-dimensional random field that maintains stochastic estimates of the occupancy state of each cell in the grid. The cell estimates are obtained by interpreting incoming range readings using probabilistic models that capture the uncertainty in the spatial information provided by the sensor. A Bayesian estimation procedure allows the incremental updating of the map using readings taken from several sensors over multiple points of view. An overview of the Occupancy Grid framework is given, and its application to a number of problems in mobile robot mapping and navigation are illustrated. It is argued that a number of robotic problem-solving activities can be performed directly on the Occupancy Grid representation. Some parallels are drawn between operations on Occupancy Grids and related image processing operations.

Elfes, Alberto↗

System Decommutes And Displays Telemetry Data

TDPlus computer program software system for decommutation of pulse-code-modulation (PCM) telemetry signals. Provides synchronization, conversion into engineering units, and display of serial bit streams. Transforms IBM PC-compatible computer into PCM-telemetry-decommutation system. Synchronizes telemetric signals data and enables conversion back into such meaningful forms as voltage, current, pressure, and the like. Also controls operation of digital-to-analog converters to ship data to paper strip charts or to parallel digital ports for offloading to other computers. Software used to process actual data only when telemetry-data-processing computer modified in accordance with specifications contained in "TDPlus TM Data Processor" (GSC-13291). Written in Turbo C and 8088 Assembly language.

Massey, D. E.↗

Scattering-induced optical polarization in thick accretion disks

A general formalism for calculating the linear polarization induced by scattering within the central funnel of a thick accretion disk is presented, and it is shown that multiple photon reflections off the funnel walls can produce polarization values of up to about 10 percent, with the polarization position angle aligned parallel to the disk symmetry axis. It is suggested that this process is responsible for the observed optical polarization levels in X-ray-selected BL Lac objects (XBLs), which generally show linear polarization percentages P less than about 10 percent. According to this interpretation, XBLs with high optical polarization are viewed at an angle of less than about 60 deg to the funnel axis and their projected polarization vectors should be preferentially aligned with the associated radio jets. The possible relevance of this model to Seyfer 1 galaxies and quasars is also discussed.

Kartje, John F.↗

Execution models for mapping programs onto distributed memory parallel computers

The problem of exploiting the parallelism available in a program to efficiently employ the resources of the target machine is addressed. The problem is discussed in the context of building a mapping compiler for a distributed memory parallel machine. The paper describes using execution models to drive the process of mapping a program in the most efficient way onto a particular machine. Through analysis of the execution models for several mapping techniques for one class of programs, we show that the selection of the best technique for a particular program instance can make a significant difference in performance. On the other hand, the results of benchmarks from an implementation of a mapping compiler show that our execution models are accurate enough to select the best mapping technique for a given program.

Sussman, Alan↗

Parallel adaptive mesh refinement techniques for plasticity problems

The accurate modeling of the nonlinear properties of materials can be computationally expensive. Parallel computing offers an attractive way for solving such problems; however, the efficient use of these systems requires the vertical integration of a number of very different software components, we explore the solution of two- and three-dimensional, small-strain plasticity problems. We consider a finite-element formulation of the problem with adaptive refinement of an unstructured mesh to accurately model plastic transition zones. We present a framework for the parallel implementation of such complex algorithms. This framework, using libraries from the SUMAA3d project, allows a user to build a parallel finite-element application without writing any parallel code. To demonstrate the effectiveness of this approach on widely varying parallel architectures, we present experimental results from an IBM SP parallel computer and an ATM-connected network of Sun UltraSparc workstations. The results detail the parallel performance of the computational phases of the application during the process while the material is incrementally loaded.

Barry, W. J.↗