Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66

Detection of edges using local geometry

Researchers described a new representation, the local geometry, for early visual processing which is motivated by results from biological vision. This representation is richer than is often used in image processing. It extracts more of the local structure available at each pixel in the image by using receptive fields that can be continuously rotated and that go to third order spatial variation. Early visual processing algorithms such as edge detectors and ridge detectors can be written in terms of various local geometries and are computationally tractable. For example, Canny's edge detector has been implemented in terms of a local geometry of order two, and a ridge detector in terms of a local geometry of order three. The edge detector in local geometry was applied to synthetic and real images and it was shown using simple interpolation schemes that sufficient information is available to locate edges with sub-pixel accuracy (to a resolution increase of at least a factor of five). This is reasonable even for noisy images because the local geometry fits a smooth surface - the Taylor series - to the discrete image data. Only local processing was used in the implementation so it can readily be implemented on parallel mesh machines such as the MPP. Researchers expect that other early visual algorithms, such as region growing, inflection point detection, and segmentation can also be implemented in terms of the local geometry and will provide sufficiently rich and robust representations for subsequent visual processing.

Gualtieri, J. A.↗

MLP: A Parallel Programming Alternative to MPI for New Shared Memory Parallel Systems

Recent developments at the NASA AMES Research Center's NAS Division have demonstrated that the new generation of NUMA based Symmetric Multi-Processing systems (SMPs), such as the Silicon Graphics Origin 2000, can successfully execute legacy vector oriented CFD production codes at sustained rates far exceeding processing rates possible on dedicated 16 CPU Cray C90 systems. This high level of performance is achieved via shared memory based Multi-Level Parallelism (MLP). This programming approach, developed at NAS and outlined below, is distinct from the message passing paradigm of MPI. It offers parallelism at both the fine and coarse grained level, with communication latencies that are approximately 50-100 times lower than typical MPI implementations on the same platform. Such latency reductions offer the promise of performance scaling to very large CPU counts. The method draws on, but is also distinct from, the newly defined OpenMP specification, which uses compiler directives to support a limited subset of multi-level parallel operations. The NAS MLP method is general, and applicable to a large class of NASA CFD codes.

Taft, James R.↗

Spatial distributions of magnetic field fluctuations in the dayside magnetosheath

In a study that tests the hypothesis that magnetosheath magnetic fields are disturbed on plasma streamlines which are connected to the quasi-parallel bow shock, magnetometer observations from the ISEE 2 and IMP 8 spacecraft are used to investigate the dayside spatial distributions of fluctuating magnetosheath fields for different interplanetary field orientations. The results suggest that although other sources such as Kelvin-Helmholtz instabilities and flux transfer processes probably contribute to the fluctuations in the magnetosheath field, the quasi-parallel shock source is an important contributor in the dayside region.

Luhmann, J. G.↗

Resource Management for Distributed Parallel Systems

Multiprocessor systems should exist in the the larger context of distributed systems, allowing multiprocessor resources to be shared by those that need them. Unfortunately, typical multiprocessor resource management techniques do not scale to large networks. The Prospero Resource Manager (PRM) is a scalable resource allocation system that supports the allocation of processing resources in large networks and multiprocessor systems. To manage resources in such distributed parallel systems, PRM employs three types of managers: system managers, job managers, and node managers. There exist multiple independent instances of each type of manager, reducing bottlenecks. The complexity of each manager is further reduced because each is designed to utilize information at an appropriate level of abstraction.

Neuman, B. Clifford↗

Parallel triangularization of substructured finite element problems

Much of the computational effort of the finite element process involves the solution of a system of linear equations. The coefficient matrix of this system, known as the global stiffness matrix, is symmetric, positive definite, and generally sparse. An important technique for reducing the time required to solve this system is substructuring or matrix partitioning. Substructuring is based on the idea of dividing a structure into pieces, each of which can then be analyzed relatively indepenently. As a result of this division, each point in the finite element discretization is either interior to a substructure or on a boundary between substructures. Contributions to the global stiffness matrix from connections between boundary points from the K(bb) matrix are reported. The triangularization of a general K(bb) matrix on a parallel machine is specifically discussed.

Leuze, M. R.↗

Performance and Application of Parallel OVERFLOW Codes on Distributed and Shared Memory Platforms

The presentation discusses recent studies on the performance of the two parallel versions of the aerodynamics CFD code, OVERFLOW_MPI and _MLP. Developed at NASA Ames, the serial version, OVERFLOW, is a multidimensional Navier-Stokes flow solver based on overset (Chimera) grid technology. The code has recently been parallelized in two ways. One is based on the explicit message-passing interface (MPI) across processors and uses the _MPI communication package. This approach is primarily suited for distributed memory systems and workstation clusters. The second, termed the multi-level parallel (MLP) method, is simple and uses shared memory for all communications. The _MLP code is suitable on distributed-shared memory systems. For both methods, the message passing takes place across the processors or processes at the advancement of each time step. This procedure is, in effect, the Chimera boundary conditions update, which is done in an explicit "Jacobi" style. In contrast, the update in the serial code is done in more of the "Gauss-Sidel" fashion. The programming efforts for the _MPI code is more complicated than for the _MLP code; the former requires modification of the outer and some inner shells of the serial code, whereas the latter focuses only on the outer shell of the code. The _MPI version offers a great deal of flexibility in distributing grid zones across a specified number of processors in order to achieve load balancing. The approach is capable of partitioning zones across multiple processors or sending each zone and/or cluster of several zones into a single processor. The message passing across the processors consists of Chimera boundary and/or an overlap of "halo" boundary points for each partitioned zone. The MLP version is a new coarse-grain parallel concept at the zonal and intra-zonal levels. A grouping strategy is used to distribute zones into several groups forming sub-processes which will run in parallel. The total volume of grid points in each group are approximately balanced. A proper number of threads are initially allocated to each group, and in subsequent iterations during the run-time, the number of threads are adjusted to achieve load balancing across the processes. Each process exploits the multitasking directives already established in Overflow.

Djomehri, M. Jahed↗

Applications of the massively parallel machine, the MasPar MP-1, to Earth sciences

The computational workload of upcoming NASA science missions, especially the ground data processing for the Earth Observing System, is projected to be quite large (in the 50 to 100 gigaFLOPS range) and corespondingly very expensive to perform using conventional supercomputer systems. High performance, general purpose massively parallel computer systems such as the MasPar MP-1 are being investigated by NASA as a more cost effective alternative. Massively parallel systems are targeted for accelerated development and maturation by NASA's upcoming five-year High Performance Computing and Communications Program. A summary of the broad range of applications currently running on the MP-1 at NASA/Goddard are presented in this paper along with descriptions of the parallel algorithmic techniques employed in five applications that have bearing on Earth sciences.

Fischer, James R.↗

Massively parallel and universal approximation of nonlinear functions using diffractive processors

Nonlinear computation is essential for a wide range of information processing tasks, yet implementing nonlinear functions using optical systems remains a challenge due to the weak and power-intensive nature of optical nonlinearities. Overcoming this limitation without relying on nonlinear optical materials could unlock unprecedented opportunities for ultrafast and parallel optical computing systems. Here, we demonstrate that large-scale nonlinear computation can be performed using linear optics through optimized diffractive processors composed of passive phase-only surfaces. In this framework, the input variables of nonlinear functions are encoded into the phase of an optical wavefront—e.g., via a spatial light modulator (SLM)—and transformed by an optimized diffractive structure with spatially varying point-spread functions to yield output intensities that approximate a large set of unique nonlinear functions–all in parallel. We provide proof establishing that this architecture serves as a universal function approximator for an arbitrary set of bandlimited nonlinear functions, also covering wavelength-multiplexed nonlinear functions as well as multi-variate and complex-valued functions that are all-optically cascadable. Our analysis also indicates the successful approximation of typical nonlinear activation functions commonly used in neural networks, including the sigmoid, tanh, ReLU (rectified linear unit), and softplus. We numerically demonstrate the parallel computation of one million distinct nonlinear functions, accurately executed at wavelength-scale spatial density at the output of a diffractive optical processor. Furthermore, we experimentally validated this framework using in situ optical learning and approximated 35 unique nonlinear functions in a single shot using a compact setup consisting of an SLM and an image sensor. These results establish diffractive optical processors as a scalable platform for massively parallel universal nonlinear function approximation, paving the way for new capabilities in analog optical computing based on linear materials.

Rahman, Md Sadman Sakib [University of California,↗

X-Ray Imagery as the Record of All Data of Interest in Hypervelocity Impact Fragment Studies

Laboratory study of hypervelocity spacecraft fragmentation has traditionally involved the collection and analysis of fragments that were caught in deceleration material surrounding the impact. This process has typically involved the disintegration of the catchment material either through chemical dissolution, or through physical excavation to recover the fragments. Due to the scale of the three impact tests—the Satellite Orbital Debris Characterization Impact Test (SOCIT), the DebriSat satellite impact test, and the DebrisLV launch vehicle impact test—the latter two using more than 12 cubic meters of polyurethane foam to capture the fragments, hese projects have used x-ray imagery to precisely locate and thus, to more efficiently extract fragments in the soft-catch material. Three years into the DebriSat fragment extraction process, a side study was initiated to explore what additional information could be discerned from the x-rays, with significant results. This study was instrumental to a rapid replacement and retooling as the project was forced to replace the x-ray system around which the extraction process had been based. The revised process continues to map the debris for extraction. The project has, in parallel, systematically addressed the limits/tolerances of what x-rays can reveal about size, shape, density, mass, velocity, energy, and deformation/damage of the fragment during the deceleration in the catchment material while replicating the original extraction mapping function. All of these features have been optimized or have sufficient understanding to characterize the basic factors that will define a complete data set extracted solely from x-ray imagery. It is an ideal time to develop such a process, with extracted fragments providing “ground truth” against image-only data, and abundant available imagery of the same fragments under both the prior and replacement x-ray technologies, which have several fundamentally different characteristics. This paper addresses the types and quality of hypervelocity fragmentation data that can be and has been extracted from x-rays. It further addresses the question of whether and under what circumstances future hypervelocity experiments can use x-ray methods to largely—or to completely—avoid the extraction process in recording all appropriate results. Lastly, this paper addresses lessons learned and how future efforts can be further optimized.

John B. Bacon↗

Direct kinematics solution architectures for industrial robot manipulators: Bit-serial versus parallel

A Very Large Scale Integration (VLSI) architecture for robot direct kinematic computation suitable for industrial robot manipulators was investigated. The Denavit-Hartenberg transformations are reviewed to exploit a proper processing element, namely an augmented CORDIC. Specifically, two distinct implementations are elaborated on, such as the bit-serial and parallel. Performance of each scheme is analyzed with respect to the time to compute one location of the end-effector of a 6-links manipulator, and the number of transistors required.

Lee, J.↗

Technology and future ground processing systems

Land-observing satellites with multiple thematic mappers will produce data at rates of 100 to 300 Mbps. When coupled with a high daily scene production rate, these rates will require new approaches to ground processing. Consideration is given here to future downlink rates and data volumes, and requirements peculiar to the future user community are discussed. The advanced technologies required to attain an operational system in the years 1985-1990 are considered, together with advances foreseen in communications, mass storage, bulk memories, and data processing. Using advanced devices, a centralized data processing system capable of handling the 100 Mbps data rate is described. New approaches, among them a parallel pipelined calibration front-end, real-time browse image production, a high bandwidth optical disk archive, regional image broadcast and massively parallel product production, are considered. A distributed system capable of handling the 300 Mbps data rate is then described. Designs for a hub system and a regional processing center are presented.

Wood, B. J.↗

NiAl alloys for structural uses

Alloys based on the intermetallic compound NiAl are of technological interest as high temperature structural alloys. These alloys possess a relatively low density, high melting temperature, good thermal conductivity, and (usually) good oxidation resistance. However, NiAl and NiAl-base alloys suffer from poor fracture resistance at low temperatures as well as inadequate creep strength at elevated temperatures. This research program explored macroalloying additions to NiAl-base alloys in order to identify possible alloying and processing routes which promote both low temperature fracture toughness and high temperature strength. Initial results from the study examined the additions of Fe, Co, and Hf on the microstructure, deformation, and fracture resistance of NiAl-based alloys. Of significance were the observations that the presence of the gamma-prime phase, based on Ni3Al, could enhance the fracture resistance if the gamma-prime were present as a continuous grain boundary film or 'necklace'; and the Ni-35Al-20Fe alloy was ductile in ribbon form despite a microstructure consisting solely of the B2 beta phase based on NiAl. The ductility inherent in the Ni-35Al-20Fe alloy was explored further in subsequent studies. Those results confirm the presence of ductility in the Ni-35Al-20Fe alloy after rapid cooling from 750 - 1000 C. However exposure at 550 C caused embrittlement; this was associated with an age-hardening reaction caused by the formation of Fe-rich precipitates. In contrast, to the Ni-35Al-20Fe alloy, exploratory research indicated that compositions in the range of Ni-35Al-12Fe retain the ordered B2 structure of NiAl, are ductile, and do not age-harden or embrittle after thermal exposure. Thus, our recent efforts have focused on the behavior of the Ni-35Al-12Fe alloy. A second parallel effort initiated in this program was to use an alternate processing technique, mechanical alloying, to improve the properties of NiAl-alloys. Mechanical alloying in the conventional sense requires ductile powder particles which, through a cold welding and fracture process, can be dispersion strengthened by submicron-sized oxide particles. Using both the Ni-35Al-Fe alloys to contain approx. 1 v/o Y2O3. Preliminary results indicate that mechanically alloyed and extruded NiAl-Fe + Y2O3 alloys when heat treated to a grain-coarsened condition, exhibit improved creep resistance at 1000 C when compared to NiAl; oxidation resistance comparable to NiAl; and fracture toughness values a factor of three better than NiAl. As a result of the research initiated on this NASA program, a subsequent project with support from Inco Alloys International is underway.

Koss, D. A.↗

Algorithms and programming tools for image processing on the MPP

Topics addressed include: data mapping and rotational algorithms for the Massively Parallel Processor (MPP); Parallel Pascal language; documentation for the Parallel Pascal Development system; and a description of the Parallel Pascal language used on the MPP.

Reeves, A. P.↗

Massively parallel processor

A brief description is given of the Massively Parallel Processor (MPP). Major applications of the MPP are in the area of image processing (where the operands are often very small integers) from very high spatial resolution passive image sensors, signal processing of radar data, and numerical modeling simulations of climate. The system can be programmed in assembly language or a high level language. Information on background, status, architecture, programming, hardware reliability, applications, and the MPP's development as a national resource for parallel algorithm research are presented in outline form.

Source record↗

Application of powder densification models to the consolidation processing of composites

Unidirectional fiber reinforced metal matrix composite tapes (containing a single layer of parallel fibers) can now be produced by plasma deposition. These tapes can be stacked and subjected to a thermomechanical treatment that results in a fully dense near net shape component. The mechanisms by which this consolidation step occurs are explored, and models to predict the effect of different thermomechanical conditions (during consolidation) upon the kinetics of densification are developed. The approach is based upon a methodology developed by Ashby and others for the simpler problem of HIP of spherical powders. The complex problem is devided into six, much simpler, subproblems, and then their predicted contributions are added to densification. The initial problem decomposition is to treat the two extreme geometries encountered (contact deformation occurring between foils and shrinkage of isolated, internal pores). Deformation of these two geometries is modelled for plastic, power law creep and diffusional flow. The results are reported in the form of a densification map.

Wadley, H. N. G.↗

Optoelectronic Tool Adds Scale Marks to Photographic Images

A simple, easy-to-use optoelectronic tool projects scale marks that become incorporated into photographic images (including film and electronic images). The sizes of objects depicted in the images can readily be measured by reference to the scale marks. The role played by the scale marks projected by this tool is the same as that of the scale marks on a ruler placed in a scene for the purpose of establishing a length scale. However, this tool offers the advantage that it can put scale marks quickly and safely in any visible location, including a location in which placement of a ruler would be difficult, unsafe, or time-consuming. The tool (see Figure 1) includes an aluminum housing, within which are mounted four laser diodes that operate at a wavelength of 670 nm. The laser diodes are spaced 1 in. (2.54 cm) apart along a baseline. The laser diodes are mounted with setscrews, which are used to adjust their beams to make them all parallel to each other and perpendicular to the baseline. During the adjustment process, the effect of the adjustments is observed by measuring the positions of the laser-beam spots on a target 80 ft (approx.24 m) away. Once the adjustments have been completed, the laser beams define three 1-in. (2.54-cm) intervals and the location of each beam is defined to within 1/16 in. (approx.1.6 mm) at any target distance out to about 80 ft (approx.24 m). The distance between the laser-beam spots as seen in an image is strictly defined only along an axis parallel to the baseline and perpendicular to the laser beam (also perpendicular to the line of sight of the camera, assuming that the camera-to-target distance is much greater than the distance between the tool and the camera lens). If a flat target surface illuminated by the laser beams is tilted with respect to the aforesaid axis, then the distance along the target surface between scale marks is proportional to the secant of the tilt angle. If one knows the tilt angle, one can correct for it. Even if one does not know the tilt angle precisely, it may not matter: For example, at a tilt of 10 , the secant is approximately 1.0154, so that the tilt error is only about 1.54 percent, which is negligibly small for a typical application in which only approximate measurements are needed.

Stevenson, Charlie↗

Vanishing dual-task interference after practice: has the bottleneck been eliminated or is it merely latent?

Practice can, in some cases, largely eliminate measured dual-task interference. Does this absence of interference indicate the absence of a processing bottleneck (defined as an inability to carry out certain stages in parallel)? The authors show that a bottleneck need not produce any observable interference, provided that there is no temporal overlap in the demand for bottleneck stages on the 2 tasks. Such a "latent" bottleneck is especially likely after practice, when central stages are short. The authors provide new evidence that a latent bottleneck occurred for a participant who produced no interference in M. Van Selst, E. Ruthruff, and J. C. Johnston (1999). These findings demonstrate that the absence of dual-task interference does not necessarily indicate the absence of a processing bottleneck.

Clinical Trial↗

Performance Assessment of OVERFLOW on Distributed Computing Environment

The aerodynamic computer code, OVERFLOW, with a multi-zone overset grid feature, has been parallelized to enhance its performance on distributed and shared memory paradigms. Practical application benchmarks have been set to assess the efficiency of code's parallelism on high-performance architectures. The code's performance has also been experimented with in the context of the distributed computing paradigm on distant computer resources using the Information Power Grid (IPG) toolkit, Globus. Two parallel versions of the code, namely OVERFLOW-MPI and -MLP, have developed around the natural coarse grained parallelism inherent in a multi-zonal domain decomposition paradigm. The algorithm invokes a strategy that forms a number of groups, each consisting of a zone, a cluster of zones and/or a partition of a large zone. Each group can be thought of as a process with one or multithreads assigned to it and that all groups run in parallel. The -MPI version of the code uses explicit message-passing based on the standard MPI library for sending and receiving interzonal boundary data across processors. The -MLP version employs no message-passing paradigm; the boundary data is transferred through the shared memory. The -MPI code is suited for both distributed and shared memory architectures, while the -MLP code can only be used on shared memory platforms. The IPG applications are implemented by the -MPI code using the Globus toolkit. While a computational task is distributed across multiple computer resources, the parallelism can be explored on each resource alone. Performance studies are achieved with some practical aerodynamic problems with complex geometries, consisting of 2.5 up to 33 million grid points and a large number of zonal blocks. The computations were executed primarily on SGI Origin 2000 multiprocessors and on the Cray T3E. OVERFLOW's IPG applications are carried out on NASA homogeneous metacomputing machines located at three sites, Ames, Langley and Glenn. Plans for the future will exploit the distributed parallel computing capability on various homogeneous and heterogeneous resources and large scale benchmarks. Alternative IPG toolkits will be used along with sophisticated zonal grouping strategies to minimize the communication time across the computer resources.

Djomehri, M. Jahed↗