Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel computers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60

Instrumentation, performance visualization, and debugging tools for multiprocessors

The need for computing power has forced a migration from serial computation on a single processor to parallel processing on multiprocessor architectures. However, without effective means to monitor (and visualize) program execution, debugging, and tuning parallel programs becomes intractably difficult as program complexity increases with the number of processors. Research on performance evaluation tools for multiprocessors is being carried out at ARC. Besides investigating new techniques for instrumenting, monitoring, and presenting the state of parallel program execution in a coherent and user-friendly manner, prototypes of software tools are being incorporated into the run-time environments of various hardware testbeds to evaluate their impact on user productivity. Our current tool set, the Ames Instrumentation Systems (AIMS), incorporates features from various software systems developed in academia and industry. The execution of FORTRAN programs on the Intel iPSC/860 can be automatically instrumented and monitored. Performance data collected in this manner can be displayed graphically on workstations supporting X-Windows. We have successfully compared various parallel algorithms for computational fluid dynamics (CFD) applications in collaboration with scientists from the Numerical Aerodynamic Simulation Systems Division. By performing these comparisons, we show that performance monitors and debuggers such as AIMS are practical and can illuminate the complex dynamics that occur within parallel programs.

Yan, Jerry C.↗

Hardware Implementation of a Bilateral Subtraction Filter

A bilateral subtraction filter has been implemented as a hardware module in the form of a field-programmable gate array (FPGA). In general, a bilateral subtraction filter is a key subsystem of a high-quality stereoscopic machine vision system that utilizes images that are large and/or dense. Bilateral subtraction filters have been implemented in software on general-purpose computers, but the processing speeds attainable in this way even on computers containing the fastest processors are insufficient for real-time applications. The present FPGA bilateral subtraction filter is intended to accelerate processing to real-time speed and to be a prototype of a link in a stereoscopic-machine- vision processing chain, now under development, that would process large and/or dense images in real time and would be implemented in an FPGA. In terms that are necessarily oversimplified for the sake of brevity, a bilateral subtraction filter is a smoothing, edge-preserving filter for suppressing low-frequency noise. The filter operation amounts to replacing the value for each pixel with a weighted average of the values of that pixel and the neighboring pixels in a predefined neighborhood or window (e.g., a 9 9 window). The filter weights depend partly on pixel values and partly on the window size. The present FPGA implementation of a bilateral subtraction filter utilizes a 9 9 window. This implementation was designed to take advantage of the ability to do many of the component computations in parallel pipelines to enable processing of image data at the rate at which they are generated. The filter can be considered to be divided into the following parts (see figure): a) An image pixel pipeline with a 9 9- pixel window generator, b) An array of processing elements; c) An adder tree; d) A smoothing-and-delaying unit; and e) A subtraction unit. After each 9 9 window is created, the affected pixel data are fed to the processing elements. Each processing element is fed the pixel value for its position in the window as well as the pixel value for the central pixel of the window. The absolute difference between these two pixel values is calculated and used as an address in a lookup table. Each processing element has a lookup table, unique for its position in the window, containing the weight coefficients for the Gaussian function for that position. The pixel value is multiplied by the weight, and the outputs of the processing element are the weight and pixel-value weight product. The products and weights are fed to the adder tree. The sum of the products and the sum of the weights are fed to the divider, which computes the sum of products the sum of weights. The output of the divider is denoted the bilateral smoothed image. The smoothing function is a simple weighted average computed over a 3 3 subwindow centered in the 9 9 window. After smoothing, the image is delayed by an additional amount of time needed to match the processing time for computing the bilateral smoothed image. The bilateral smoothed image is then subtracted from the 3 3 smoothed image to produce the final output. The prototype filter as implemented in a commercially available FPGA processes one pixel per clock cycle. Operation at a clock speed of 66 MHz has been demonstrated, and results of a static timing analysis have been interpreted as suggesting that the clock speed could be increased to as much as 100 MHz.

Huertas, Andres↗

Simplifying Geospatial Workflows with GeoGridFusion

Growing demands to understand PV deployment and reliability across an expanding range of climates, along with increasing computational power, are driving the need for simplified computing tools that support large-scale analysis and geospatial workflows. While existing open-source libraries and PV system modeling tools offer extensive collections of empirical and physical models, they often lack the ability to scale to large geospatial datasets. The tool presented here, GeoGridFusion, enables PV modelers to store and utilize gridded geospatial data from sources such as NSRDB, PVGIS, and others. GeoGridFusion harmonizes diverse datasets, taxonomies, and nomenclatures, and includes utilities that support intuitive geospatial area selection.

14 SOLAR ENERGY↗

A Mixed Integer Efficient Global Optimization Algorithm with Multiple Infill Strategy - Applied to a Wing Topology Optimization Problem

With the advancement in high performance computing and numerical optimization techniques,engineering design optimization problems are becoming more complex, larger scale,higher fidelity, and computationally more demanding, requiring longer run times than ever before. There exists methodologies and techniques that can address some of these challenges but very few can address all, and most are limited in the extent that these concerns can be addressed. With the goal of addressing such challenging engineering problems, we developed anew optimization framework, named AMIEGO, that combines concepts from surrogate-based optimization approaches, gradient-based numerical methods, Partial Least Squares, evolutionary algorithms, and Branch-and-Bound, providing newer capabilities that were not previouslyperceived. However, the original version of this framework, in the process of adaptive samplingto explore and exploit the design space, finds only a single sample point per iteration. The efforthere builds upon this previously developed optimization framework to include multiple infillsampling capability that combines the concept of generalized expected improvement function,unsupervised learning, and multi-objective evolutionary technique. To demonstrate, AMIEGOwith the multiple infill capability (called AMIEGO-MIMOS) solves a series of increasingly difficultengineering design optimization problems. The results reveal the performance of the newapproach is problem dependent. When applied to a ten-bar truss problem, the newly proposedmultiple infill strategy consistently leads to a better design solutions when compared to theexisting CPTV method (implemented with the context of the AMIEGO framework). On theother hand, when applied to a mixed-integer high fidelity wing topology optimization problem- MIMOS, despite showing a steeper convergence at the start, eventually leads to an inferiorsolution as compared to CPTV approach. These results also reveal that a small number ofstarting points, in general, are sufficient to lead to a good overall solution.

Mixed-integer optimization↗

A parallel finite-difference method for computational aerodynamics

A finite-difference scheme for solving complex three-dimensional aerodynamic flow on parallel-processing supercomputers is presented. The method consists of a basic flow solver with multigrid convergence acceleration, embedded grid refinements, and a zonal equation scheme. Multitasking and vectorization have been incorporated into the algorithm. Results obtained include multiprocessed flow simulations from the Cray X-MP and Cray-2. Speedups as high as 3.3 for the two-dimensional case and 3.5 for segments of the three-dimensional case have been achieved on the Cray-2. The entire solver attained a factor of 2.7 improvement over its unitasked version on the Cray-2. The performance of the parallel algorithm on each machine is analyzed.

Swisshelm, Julie M.↗

The CP-PAW Code Package for First-Principles Calculations from a User’s Perspective

CP-PAW is a combined electronic structure and ab initio molecular dynamics code to perform mixed quantum and classical simulations of atomistic condensed phase systems, such as solids, liquids, and molecular systems. As the name suggests, the CP-PAW code unifies the all-electron projector augmented-wave (PAW) method with the Car–Parrinello (CP) approach to determine not only the electronic and nuclear ground states of condensed matter but also to study their properties and dynamics. In addition to briefly outlining the underlying theory, the focus will be on the unique aspects of CP-PAW and how to correctly employ them as a user. How to install CP-PAW using the new build system will also be briefly mentioned.

Blöchl, Peter E [Institute for Theoretical Physic↗

Towards Seamless Interoperability of MPI-OpenMP Applications

A chasm exists between mathematical software libraries written for MPI-based applications and those written for OpenMP applications. Recently, however, PETSc enables the simple use of its MPI-based linear solvers from OpenMP applications. Separately, the MPICH MPI development team has started a new project to allow almost seamless MPI use in OpenMP applications. Both proposed approaches would result in a similar user experience. Here, we discuss the reasons for these projects and their potential for providing more numerical library choices for OpenMP applications, including the unlimited assortment of linear solvers available in PETSc. In addition, we present the performance of an application using the first approach, demonstrating its efficacy.

MPI↗

Mirroring within the Fokker-Planck formulation of cosmic ray pitch angle scattering in homogeneous magnetic turbulence

The Fokker-Planck coefficient for pitch angle scattering, appropriate for cosmic rays in homogeneous, stationary, magnetic turbulence, is computed from first principles. No assumptions are made concerning any special statistical symmetries the random field may have. This result can be used to compute the parallel diffusion coefficient for high energy cosmic rays moving in strong turbulence, or low energy cosmic rays moving in weak turbulence. Becuase of the generality of the magnetic turbulence which is allowed in this calculation, special interplanetary magnetic field features such as discontinuities, or particular wave modes, can be included rigorously. The reduction of this results to previously available expressions for the pitch angle scattering coefficient in random field models with special symmetries is discussed. The general existance of a Dirac delta function in the pitch angle scattering coefficient is demonstrated. It is proved that this delta function is the Fokker-Planck prediction for pitch angle scattering due to mirroring in the magnetic field.

Goldstein, M. L.↗

A 30-day forecast experiment with the GISS model and updated sea surface temperatures

The GISS model was used to compute two parallel global 30-day forecasts for the month January 1974. In one forecast, climatological January sea surface temperatures were used, while in the other observed sea temperatures were inserted and updated daily. A comparison of the two forecasts indicated no clear-cut beneficial effect of daily updating of sea surface temperatures. Despite the rapid decay of daily predictability, the model produced a 30-day mean forecast for January 1974 that was generally superior to persistence and climatology when evaluated over either the globe or the Northern Hemisphere, but not over smaller regions.

Spar, J.↗

Mirroring in the Fokker-Planck coefficient for cosmic-ray pitch-angle scattering in homogeneous magnetic turbulence

The Fokker-Planck coefficient for pitch-angle scattering, appropriate for cosmic rays in homogeneous stationary magnetic turbulence is computed without making any specific assumptions concerning the statistical symmetries of the random field. The Fokker-Planck coefficient obtained can be used to compute the parallel diffusion coefficient for high-energy cosmic rays propagating in the presence of strong turbulence, or for low-energy cosmic rays in the presence of weak turbulence. Because of the generality of magnetic turbulence allowed for in the analysis, special interplanetary magnetic field features, such as discontinuities or particular wave modes, can be included rigorously.

Goldstein, M. L.↗

Basic cluster compression algorithm

Feature extraction and data compression of LANDSAT data is accomplished by BCCA program which reduces costs associated with transmitting, storing, distributing, and interpreting multispectral image data. Algorithm uses spatially local clustering to extract features from image data to describe spectral characteristics of data set. Approach requires only simple repetitive computations, and parallel processing can be used for very high data rates. Program is written in FORTRAN IV for batch execution and has been implemented on SEL 32/55.

Hilbert, E. E.↗

Application of the multigrid method to grid generation

The multigrid method (MGM), used to numerically solve the pair of nonlinear elliptic equations commonly used to generate two dimensional boundary-fitted coordinate systems is discussed. Two different geometries are considered: one involving a coordinate system fitted about a circle and the other selected for an impinging jet flow problem. Two different relaxation schemes are tried: one is successive point overrelaxation and the other is a four-color scheme vectorizeable to take advantage of a parallel processor computer for greater computational speed. Results using MGM are compared with those using SOR (doing successive overrelaxations with the corresponding relaxation scheme on the fine grid only). It is found that MGM becomes significantly more effective than SOR as more accuracy is demanded and as more corrective grids, or more grid points, are used. For the accuracy required, it is found that MGM is two to three times faster than SOR in computing time. With the four-color relaxation scheme as applied to the impinging jet problem, the advantage of MGM over SOR is not as great. This may be due to the effect of a poor initial guess on MGM for this problem.

Ohring, S.↗

A linear decomposition method for large optimization problems. Blueprint for development

A method is proposed for decomposing large optimization problems encountered in the design of engineering systems such as an aircraft into a number of smaller subproblems. The decomposition is achieved by organizing the problem and the subordinated subproblems in a tree hierarchy and optimizing each subsystem separately. Coupling of the subproblems is accounted for by subsequent optimization of the entire system based on sensitivities of the suboptimization problem solutions at each level of the tree to variables of the next higher level. A formalization of the procedure suitable for computer implementation is developed and the state of readiness of the implementation building blocks is reviewed showing that the ingredients for the development are on the shelf. The decomposition method is also shown to be compatible with the natural human organization of the design process of engineering systems. The method is also examined with respect to the trends in computer hardware and software progress to point out that its efficiency can be amplified by network computing using parallel processors.

Sobieszczanski-Sobieski, J.↗

Aerothermal loads analysis for high speed flow over a quilted surface configuration

Attention is given to hypersonic laminar flow over a quilted surface configuration that simulates an array of Space Shuttle Thermal Protection System panels bowed in a spherical shape as a result of thermal gradient through the panel thickness. Pressure and heating loads to the surface are determined. The flow field over the configuration was mathematically modeled by means of time-dependent, three-dimensional conservation of mass, momentum, and energy equations. A boundary mapping technique was then used to obtain a rectangular, parallel piped computational domain, and an explicit MacCormack (1972) explicit time-split predictor corrector finite difference algorithm was used to obtain steady state solutions. Total integrated heating loads vary linearly with bowed height when this value does not exceed the local boundary layer thickness.

Olsen, G. C.↗

Parallel structures in human and computer memory

If one thinks of our experiences as being recorded continuously on film, then human memory can be compared to a film library that is indexed by the contents of the film strips stored in it. Moreover, approximate retrieval cues suffice to retrieve information stored in this library. One recognizes a familiar person in a fuzzy photograph or a familiar tune played on a strange instrument. A computer memory that would allow a computer to recognize patterns and to recall sequences the way humans do is constructed. Such a memory is remarkably similiar in structure to a conventional computer memory and also to the neural circuits in the cortex of the cerebellum of the human brain. It is concluded that the frame problem of artificial intelligence could be solved by the use of such a memory if one were able to encode information about the world properly.

Kanerva, P.↗

A research program in advanced information systems

Topics addressed cover: identification of large space structure dynamics; special-purpose architectures for computational fluid dynamics; fault-tolerant processor architectures; data flow techniques; software specification and verification tools; management of software development; software environment for concurrent computing; and parallel algorithms and architectures for the solution of partial differential equations.

Vandervelde, Wallace E.↗

Numerical simulations of aerodynamic contribution of flows about a space-plane-type configuration

The slightly supersonic viscous flow about the space-plane under development at the National Aerospace Laboratory (NAL) in Japan was simulated numerically using the LU-ADI algorithm. The wind-tunnel testing for the same plane also was conducted with the computations in parallel. The main purpose of the simulation is to capture the phenomena which have a great deal of influence to the aerodynamic force and efficiency but is difficult to capture by experiments. It includes more accurate representation of vortical flows with high angles of attack of an aircraft. The space-plane shape geometry simulated is the simplified model of the real space-plane, which is a combination of a flat and slender body and a double-delta type wing. The comparison between experimental results and numerical ones will be done in the near future. It could be said that numerical results show the qualitatively reliable phenomena.

Matsushima, Kisa↗

Optoelectronic analogs of self-programming neural nets - Architecture and methodologies for implementing fast stochastic learning by simulated annealing

Self-organization and learning is a distinctive feature of neural nets and processors that sets them apart from conventional approaches to signal processing. It leads to self-programmability which alleviates the problem of programming complexity in artificial neural nets. In this paper architectures for partitioning an optoelectronic analog of a neural net into distinct layers with prescribed interconnectivity pattern to enable stochastic learning by simulated annealing in the context of a Boltzmann machine are presented. Stochastic learning is of interest because of its relevance to the role of noise in biological neural nets. Practical considerations and methodologies for appreciably accelerating stochastic learning in such a multilayered net are described. These include the use of parallel optical computing of the global energy of the net, the use of fast nonvolatile programmable spatial light modulators to realize fast plasticity, optical generation of random number arrays, and an adaptive noisy thresholding scheme that also makes stochastic learning more biologically plausible. The findings reported predict optoelectronic chips that can be used in the realization of optical learning machines.

Farhat, Nabil H.↗