Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44

A parallel Householder tridiagonalization stratagem using scattered square decomposition

The parallel stratagem in this paper uses scattered square decomposition, introduced by Fox (1985), for its data assignment and then exploits parallelism in the solution steps of the sequential Householder tridiagonalization algorithm. One may condense a real symmetric full matrix A of order n into a tridiagonal form by the stratagem in concurrent machines where N(=D-squared) processors are used. Expressions for efficiency and speedup are given for the evaluation of the stratagem. An alternative stratagem which requires less data transmission but more computations is also discussed. The results shown that the Householder method of tridiagonalization may be implemented on a concurrent machine efficiently by scattered square decomposition provided that the number of matrix elements contained in each processor is much larger than the number of processors of the concurrent machine, and the ratio of the time to transmit one data item from one processor to any other processor to the time to perform a floating-point arithmetic operation is small enough.

Chang, H. Y.↗

Introduction to a system for implementing neural net connections on SIMD architectures

Neural networks have attracted much interest recently, and using parallel architectures to simulate neural networks is a natural and necessary application. The SIMD model of parallel computation is chosen, because systems of this type can be built with large numbers of processing elements. However, such systems are not naturally suited to generalized elements. A method is proposed that allows an implementation of neural network connections on massively parallel SIMD architectures. The key to this system is an algorithm permitting the formation of arbitrary connections between the neurons. A feature is the ability to add new connections quickly. It also has error recovery ability and is robust over a variety of network topologies. Simulations of the general connection system, and its implementation on the Connection Machine, indicate that the time and space requirements are proportional to the product of the average number of connections per neuron and the diameter of the interconnection network.

Tomboulian, Sherryl↗

Optimizing Input/Output Using Adaptive File System Policies

Parallel input/output characterization studies and experiments with flexible resource management algorithms indicate that adaptivity is crucial to file system performance. In this paper we propose an automatic technique for selecting and refining file system policies based on application access patterns and execution environment. An automatic classification framework allows the file system to select appropriate caching and pre-fetching policies, while performance sensors provide feedback used to tune policy parameters for specific system environments. To illustrate the potential performance improvements possible using adaptive file system policies, we present results from experiments involving classification-based and performance-based steering.

Madhyastha, Tara M.↗

Balancing Contention and Synchronization on the Intel Paragon

The Intel Paragon is a mesh-connected distributed memory parallel computer. It uses an oblivious and deterministic message routing algorithm: this permits us to develop highly optimized schedules for frequently needed communication patterns. The complete exchange is one such pattern. Several approaches are available for carrying it out on the mesh. We study an algorithm developed by Scott. This algorithm assumes that a communication link can carry one message at a time and that a node can only transmit one message at a time. It requires global synchronization to enforce a schedule of transmissions. Unfortunately global synchronization has substantial overhead on the Paragon. At the same time the powerful interconnection mechanism of this machine permits 2 or 3 messages to share a communication link with minor overhead. It can also overlap multiple message transmission from the same node to some extent. We develop a generalization of Scott's algorithm that executes complete exchange with a prescribed contention. Schedules that incur greater contention require fewer synchronization steps. This permits us to tradeoff contention against synchronization overhead. We describe the performance of this algorithm and compare it with Scott's original algorithm as well as with a naive algorithm that does not take interconnection structure into account. The Bounded contention algorithm is always better than Scott's algorithm and outperforms the naive algorithm for all but the smallest message sizes. The naive algorithm fails to work on meshes larger than 12 x 12. These results show that due consideration of processor interconnect and machine performance parameters is necessary to obtain peak performance from the Paragon and its successor mesh machines.

Bokhari, Shahid H.↗

Genetically Engineered Microelectronic Infrared Filters

A genetic algorithm is used for design of infrared filters and in the understanding of the material structure of a resonant tunneling diode. These two components are examples of microdevices and nanodevices that can be numerically simulated using fundamental mathematical and physical models. Because the number of parameters that can be used in the design of one of these devices is large, and because experimental exploration of the design space is unfeasible, reliable software models integrated with global optimization methods are examined The genetic algorithm and engineering design codes have been implemented on massively parallel computers to exploit their high performance. Design results are presented for the infrared filter showing new and optimized device design. Results for nanodevices are presented in a companion paper at this workshop.

Cwik, Tom↗

Perspectives on the Future of CFD

This viewgraph presentation gives an overview of the future of computational fluid dynamics (CFD), which in the past has pioneered the field of flow simulation. Over time CFD has progressed as computing power. Numerical methods have been advanced as CPU and memory capacity increases. Complex configurations are routinely computed now and direct numerical simulations (DNS) and large eddy simulations (LES) are used to study turbulence. As the computing resources changed to parallel and distributed platforms, computer science aspects such as scalability (algorithmic and implementation) and portability and transparent codings have advanced. Examples of potential future (or current) challenges include risk assessment, limitations of the heuristic model, and the development of CFD and information technology (IT) tools.

Kwak, Dochan↗

Molecular Simulations in Astrobiology

One of the main goals of astrobiology is to understand the origin of cellular life. In the absence of any record of the earliest ancestors of contemporary cells, protocells, the most direct way to test our understanding of their characteristics is to construct laboratory models of protocells. Such efforts, currently underway in the NASA Astrobiology Program, are accompanied by computational studies aimed at explaining self-organization of simple molecules into ordered structures and developing designs of molecules that are capable of performing protocellular functions. Many of these functions, such as importing nutrients, capturing and storing energy, and responding to changes in the environment, are carried out by proteins bound to membranes. We use computer simulations to address the following, questions about these proteins: (1) How do small proteins (peptides) organize themselves into ordered structures at water-membrane interfaces and insert into membranes? (2) How do peptides aggregate to form membrane-spannin(y structures (e.g., channels)? (3) By what mechanisms do such aggregates perform their functions? The simulations are performed using the molecular dynamics (MD) method. In this method, Newton's equations of motion for each atom in the system are solved iteratively. At each time step, the forces exerted on each atom by the remaining atoms are evaluated by dividing them into two parts. Short-range forces are calculated directly in real space while long-range forces are evaluated in reciprocal space, usually using a particle-mesh algorithm which is of order O(NlnN). Currently, a time step of 2 femtoseconds is typically used, thereby making studies of problems occurring on multi-nanosecond time scales (10(exp 6) - 10(exp 8) time steps) accessible. To address a broader range of problems, simulations need to be extended by three orders of magnitude. Such an extension requires both algorithmic improvements and codes scalable to a large number of parallel processors. Work in this direction is in progress. Two specific series of simulations that demonstrate how peptides self-organize and function in membranes are discussed. In one series of simulations, it was shown that nonpolar peptides, disordered in water, translocate to the nonpolar interior of the membrane and, simultaneously, fold into two different helical structures, which remain in equilibrium. Once in the membrane, the peptides can readily change their orientation, especially in response to local electric fields. This structural and orientational flexibility of peptides with changing conditions may have provided a mechanism of transmitting signals between the environment and the interior of the protocell. In another series of simulations, the mechanism by which a simple protein channel efficiently mediates proton transport across membranes was investigated. This process is a key step in cellular bioenergetics. In the channel under study, proton transport is gated by four histidines that occlude the channel pore. The simulations demonstrate that protons move through the gate by a "shuttle" mechanism, wherein one histidine is protonated on the extracellular side and, subsequently, the proton bound on the opposite side is released.

Pohorille, Andrew↗

Toward GEOS-6, A Global Cloud System Resolving Atmospheric Model

NASA is committed to observing and understanding the weather and climate of our home planet through the use of multi-scale modeling systems and space-based observations. Global climate models have evolved to take advantage of the influx of multi- and many-core computing technologies and the availability of large clusters of multi-core microprocessors. GEOS-6 is a next-generation cloud system resolving atmospheric model that will place NASA at the forefront of scientific exploration of our atmosphere and climate. Model simulations with GEOS-6 will produce a realistic representation of our atmosphere on the scale of typical satellite observations, bringing a visual comprehension of model results to a new level among the climate enthusiasts. In preparation for GEOS-6, the agency's flagship Earth System Modeling Framework [JDl] has been enhanced to support cutting-edge high-resolution global climate and weather simulations. Improvements include a cubed-sphere grid that exposes parallelism; a non-hydrostatic finite volume dynamical core, and algorithm designed for co-processor technologies, among others. GEOS-6 represents a fundamental advancement in the capability of global Earth system models. The ability to directly compare global simulations at the resolution of spaceborne satellite images will lead to algorithm improvements and better utilization of space-based observations within the GOES data assimilation system

Putman, William M.↗

Real-time dynamics and control strategies for space operations of flexible structures

This project (NAG9-574) was meant to be a three-year research project. However, due to NASA's reorganizations during 1992, the project was funded only for one year. Accordingly, every effort was made to make the present final report as if the project was meant to be for one-year duration. Originally, during the first year we were planning to accomplish the following: we were to start with a three dimensional flexible manipulator beam with articulated joints and with a linear control-based controller applied at the joints; using this simple example, we were to design the software systems requirements for real-time processing, introduce the streamlining of various computational algorithms, perform the necessary reorganization of the partitioned simulation procedures, and assess the potential speed-up realization of the solution process by parallel computations. The three reports included as part of the final report address: the streamlining of various computational algorithms; the necessary reorganization of the partitioned simulation procedures, in particular the observer models; and an initial attempt of reconfiguring the flexible space structures.

Park, K. C.↗

Bit-parallel arithmetic in a massively-parallel associative processor

A simple but powerful new architecture based on a classical associative processor model is presented. Algorithms for performing the four basic arithmetic operations both for integer and floating point operands are described. For m-bit operands, the proposed architecture makes it possible to execute complex operations in O(m) cycles as opposed to O(m exp 2) for bit-serial machines. A word-parallel, bit-parallel, massively-parallel computing system can be constructed using this architecture with VLSI technology. The operation of this system is demonstrated for the fast Fourier transform and matrix multiplication.

Scherson, Isaac D.↗

Parallel Climate Data Assimilation PSAS Package

We have designed and implemented a set of highly efficient and highly scalable algorithms for an unstructured computational package, the PSAS data assimilation package, as demonstrated by detailed performance analysis of systematic runs on up to 512node Intel Paragon. The equation solver achieves a sustained 18 Gflops performance. As the results, we achieved an unprecedented 100-fold solution time reduction on the Intel Paragon parallel platform over the Cray C90. This not only meets and exceeds the DAO time requirements, but also significantly enlarges the window of exploration in climate data assimilations.

PSAS data scalable algorithms Intel Paragon 512nod↗

Interpreting Broad Double-Peaked Emission Lines in Active Galactic Nuclei

The principal objectives of this project were to probe the inner regions of active galactic nuclei and to test general relativity in the strong-field limit. The approach takes advantage of broad atomic line emission observed from material deep in the potential well of an active galactic nucleus which contains key information as to the physics of the system. Line profiles in a wide range of wavebands from optical to X-ray have provided compelling evidence of the existence of a relativistic accretion disk around a supermassive black hole in a number of galaxies. The simplest model posits a geometrically thin disk in Keplerian orbit, with general relativistic effects in evidence. This model is the point of departure for the proposed work. We developed a high-performance numerical code to calculate photon trajectories in a Schwarzschild or Kerr metric and implemented it on parallel supercomputers. This code includes a general purpose ray tracer that calculates line profiles, light curves, and other observable quantities for a wide variety of emitter configurations. The versatility comes from the fact that the ray tracing algorithm does not depend on any symmetries regarding emitter locations. The speed comes from parallel implementation which enables us to sample hitherto unattainable volumes of disk model parameter space. During the period 1 March 1997 through 28 February 1998, two papers, supported in whole or in part by this grant, were published in refereed journals. They are reproduced in their entirety in the next two sections of this report.

Halpern, Jules↗

Semi-automatic process partitioning for parallel computation

On current multiprocessor architectures one must carefully distribute data in memory in order to achieve high performance. Process partitioning is the operation of rewriting an algorithm as a collection of tasks, each operating primarily on its own portion of the data, to carry out the computation in parallel. A semi-automatic approach to process partitioning is considered in which the compiler, guided by advice from the user, automatically transforms programs into such an interacting task system. This approach is illustrated with a picture processing example written in BLAZE, which is transformed into a task system maximizing locality of memory reference.

Koelbel, Charles↗

A high-performance FFT algorithm for vector supercomputers

Many traditional algorithms for computing the fast Fourier transform (FFT) on conventional computers are unacceptable for advanced vector and parallel computers because they involve nonunit, power-of-two memory strides. A practical technique for computing the FFT that avoids all such strides and appears to be near-optimal for a variety of current vector and parallel computers is presented. Performance results of a program based on this technique are given. Notable among these results is that a FORTRAN implementation of this algorithm on the CRAY-2 runs up to 77-percent faster than Cray's assembly-coded library routine.

Bailey, David H.↗

An Evaluation of the Measurement Requirements for an In-Situ Wake Vortex Detection System

Results of a numerical simulation are presented to determine the feasibility of estimating the location and strength of a wake vortex from imperfect in-situ measurements. These estimates could be used to provide information to a pilot on how to avoid a hazardous wake vortex encounter. An iterative algorithm based on the method of secants was used to solve the four simultaneous equations describing the two-dimensional flow field around a pair of parallel counter-rotating vortices of equal and constant strength. The flow field information used by the algorithm could be derived from measurements from flow angle sensors mounted on the wing-tip of the detecting aircraft and an inertial navigation system. The study determined the propagated errors in the estimated location and strength of the vortex which resulted from random errors added to theoretically perfect measurements. The results are summarized in a series of charts and a table which make it possible to estimate these propagated errors for many practical situations. The situations include several generator-detector airplane combinations, different distances between the vortex and the detector airplane, as well as different levels of total measurement error.

Fuhrmann, Henri D.↗

Towards a Multisensor Approach to Improve on Current TRMM Retrievals of Clouds and Precipitation

The Tropical Rainfall Measuring Mission (TRMM) was designed to measure tropical rainfall and its variation from a low inclination orbiting satellite. The TRMM payload was carefully chosen to overcome a number of limitations of past satellite observing systems. This payload is predicated on the combination of active and passive observations from the TRMM Precipitation Radar (PR) and TRMM Microwave Imager (TMI) and Visible and Infrared Scanner (VIRS). Our research over the past three years has been devoted to the challenge of developing the most effective way of combining complementary information from these sensors to provide the most consistent estimate of precipitation. We have approached this problem from three directions. The first was to carry out preliminary analysis of passive microwave and infrared data from the TMI and VIRS instruments to understand the character of clear and cloudy skies in the basis defined by polarization and brightness temperature differences. Using this information as a foundation, the properties of two retrieval algorithms were analyzed, one for retrieving ice clouds from VIRS that was developed in parallel with this project and the other for rainfall from the TMI. Finally, the knowledge gleaned from each of these studies, coupled with ancillary data from NWP models and a broadband radiative transfer model, was used to create and algorithm for synthesizing the principal components of the Earth's energy budget from the basic building blocks of the atmosphere, gases, clouds, and precipitation. Principal results from each of these areas of research and their role in the TRMM and climate communities are summarized.

Stephens, Graeme L.↗

Supercomputing on massively parallel bit-serial architectures

Research on the Goodyear Massively Parallel Processor (MPP) suggests that high-level parallel languages are practical and can be designed with powerful new semantics that allow algorithms to be efficiently mapped to the real machines. For the MPP these semantics include parallel/associative array selection for both dense and sparse matrices, variable precision arithmetic to trade accuracy for speed, micro-pipelined train broadcast, and conditional branching at the processing element (PE) control unit level. The preliminary design of a FORTRAN-like parallel language for the MPP has been completed and is being used to write programs to perform sparse matrix array selection, min/max search, matrix multiplication, Gaussian elimination on single bit arrays and other generic algorithms. A description is given of the MPP design. Features of the system and its operation are illustrated in the form of charts and diagrams.

Iobst, Ken↗

Parallel Climate Data Assimilation PSAS Package Achieves 18 GFLOPs on 512-Node Intel Paragon

Several algorithms were added to the Physical-space Statistical Analysis System (PSAS) from Goddard, which assimilates observational weather data by correcting for different levels of uncertainty about the data and different locations for mobile observation platforms. The new algorithms and use of the 512-node Intel Paragon allowed a hundred-fold decrease in processing time.

weather prediction climate modeling data assimilat↗