Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56

An interactive parallel programming environment applied in atmospheric science

This article introduces an interactive parallel programming environment (IPPE) that simplifies the generation and execution of parallel programs. One of the tasks of the environment is to generate message-passing parallel programs for homogeneous and heterogeneous computing platforms. The parallel programs are represented by using visual objects. This is accomplished with the help of a graphical programming editor that is implemented in Java and enables portability to a wide variety of computer platforms. In contrast to other graphical programming systems, reusable parts of the programs can be stored in a program library to support rapid prototyping. In addition, runtime performance data on different computing platforms is collected in a database. A selection process determines dynamically the software and the hardware platform to be used to solve the problem in minimal wall-clock time. The environment is currently being tested on a Grand Challenge problem, the NASA four-dimensional data assimilation system.

Interactive Display Devices↗

Planning in subsumption architectures

A subsumption planner using a parallel distributed computational paradigm based on the subsumption architecture for control of real-world capable robots is described. Virtual sensor state space is used as a planning tool to visualize the robot's anticipated effect on its environment. Decision sequences are generated based on the environmental situation expected at the time the robot must commit to a decision. Between decision points, the robot performs in a preprogrammed manner. A rudimentary, domain-specific partial world model contains enough information to extrapolate the end results of the rote behavior between decision points. A collective network of predictors operates in parallel with the reactive network forming a recurrrent network which generates plans as a hierarchy. Details of a plan segment are generated only when its execution is imminent. The use of the subsumption planner is demonstrated by a simple maze navigation problem.

Chalfant, Eugene C.↗

Injector Design Tool Improvements: User's manual for FDNS V.4.5

The major emphasis of the current effort is in the development and validation of an efficient parallel machine computational model, based on the FDNS code, to analyze the fluid dynamics of a wide variety of liquid jet configurations for general liquid rocket engine injection system applications. This model includes physical models for droplet atomization, breakup/coalescence, evaporation, turbulence mixing and gas-phase combustion. Benchmark validation cases for liquid rocket engine chamber combustion conditions will be performed for model validation purpose. Test cases may include shear coaxial, swirl coaxial and impinging injection systems with combinations LOXIH2 or LOXISP-1 propellant injector elements used in rocket engine designs. As a final goal of this project, a well tested parallel CFD performance methodology together with a user's operation description in a final technical report will be reported at the end of the proposed research effort.

Chen, Yen-Sen↗

Fluid Management of and Flame Spread Across Liquid Pools

The goal of our research on flame spread across pools of liquid fuel remains the quantitative identification of the mechanisms that control the rate and nature of flame spread when the initial temperature of the liquid pool is below the fuel's flash point temperature. As described in, four microgravity (mu-g) sounding rocket flights examined the effect of forced opposed airflow over a 2.5 cm deep x 2 cm wide x 30 cm long pool of 1-butanol. Among many unexpected findings, it was observed that the flame spread is much slower and steadier than in 1g where flame spread has a pulsating character. Our numerical model, restricted to two dimensions, had predicted faster, pulsating flame spread in mu-g. In a test designed to achieve a more 2-D experiment, our investigation of a shallow, wide pool (2 mm deep x 78 mm wide x 30 cm long) was unsuccessful in mu-g, due to an unexpectedly long time required to fill the tray. As such, the most recent Spread Across Liquids (SAL) sounding rocket experiment had two principal objectives: 1) determine if pulsating flame spread in deep fuel trays would occur under the conditions that a state-of-the-art computational combustion code and short-duration drop tower tests predict; and 2) determine if a long, rectangular, shallow fuel tray could achieve a visibly flat liquid surface across the whole tray without spillage in the mu-g time allotted. If the second objective was met, the shallow tray was to be ignited to determine the nature of flame spread in mu-g for this geometry. For the first time in the experiment series, two fuel trays - one deep (30 cm long x 2 cm wide x 25 mm deep) and one shallow (same length and width, but 2 mm deep)-- were flown. By doing two independent experiments in a single flight, a significant cost savings was realized. In parallel, the computational objective was to modify the code to improve agreement with earlier results. This last objective was achieved by modifying the fuel mass diffusivity and adding a parameter to correct for radiative and lateral heat loss.

Ross, H. D.↗

A parallel-vector algorithm for rapid structural analysis on high-performance computers

A fast, accurate Choleski method for the solution of symmetric systems of linear equations is presented. This direct method is based on a variable-band storage scheme and takes advantage of column heights to reduce the number of operations in the Choleski factorization. The method employs parallel computation in the outermost DO-loop and vector computation via the 'loop unrolling' technique in the innermost DO-loop. The method avoids computations with zeros outside the column heights, and as an option, zeros inside the band. The close relationship between Choleski and Gauss elimination methods is examined. The minor changes required to convert the Choleski code to a Gauss code to solve non-positive-definite symmetric systems of equations are identified. The results for two large-scale structural analyses performed on supercomputers, demonstrate the accuracy and speed of the method.

Storaasli, Olaf O.↗

A parallel-vector algorithm for rapid structural analysis on high-performance computers

A fast, accurate Choleski method for the solution of symmetric systems of linear equations is presented. This direct method is based on a variable-band storage scheme and takes advantage of column heights to reduce the number of operations in the Choleski factorization. The method employs parallel computation in the outermost DO-loop and vector computation via the loop unrolling technique in the innermost DO-loop. The method avoids computations with zeros outside the column heights, and as an option, zeros inside the band. The close relationship between Choleski and Gauss elimination methods is examined. The minor changes required to convert the Choleski code to a Gauss code to solve non-positive-definite symmetric systems of equations are identified. The results for two large scale structural analyses performed on supercomputers, demonstrate the accuracy and speed of the method.

Storaasli, Olaf O.↗

Instrumentation, performance visualization, and debugging tools for multiprocessors

The need for computing power has forced a migration from serial computation on a single processor to parallel processing on multiprocessor architectures. However, without effective means to monitor (and visualize) program execution, debugging, and tuning parallel programs becomes intractably difficult as program complexity increases with the number of processors. Research on performance evaluation tools for multiprocessors is being carried out at ARC. Besides investigating new techniques for instrumenting, monitoring, and presenting the state of parallel program execution in a coherent and user-friendly manner, prototypes of software tools are being incorporated into the run-time environments of various hardware testbeds to evaluate their impact on user productivity. Our current tool set, the Ames Instrumentation Systems (AIMS), incorporates features from various software systems developed in academia and industry. The execution of FORTRAN programs on the Intel iPSC/860 can be automatically instrumented and monitored. Performance data collected in this manner can be displayed graphically on workstations supporting X-Windows. We have successfully compared various parallel algorithms for computational fluid dynamics (CFD) applications in collaboration with scientists from the Numerical Aerodynamic Simulation Systems Division. By performing these comparisons, we show that performance monitors and debuggers such as AIMS are practical and can illuminate the complex dynamics that occur within parallel programs.

Yan, Jerry C.↗

Hardware Implementation of a Bilateral Subtraction Filter

A bilateral subtraction filter has been implemented as a hardware module in the form of a field-programmable gate array (FPGA). In general, a bilateral subtraction filter is a key subsystem of a high-quality stereoscopic machine vision system that utilizes images that are large and/or dense. Bilateral subtraction filters have been implemented in software on general-purpose computers, but the processing speeds attainable in this way even on computers containing the fastest processors are insufficient for real-time applications. The present FPGA bilateral subtraction filter is intended to accelerate processing to real-time speed and to be a prototype of a link in a stereoscopic-machine- vision processing chain, now under development, that would process large and/or dense images in real time and would be implemented in an FPGA. In terms that are necessarily oversimplified for the sake of brevity, a bilateral subtraction filter is a smoothing, edge-preserving filter for suppressing low-frequency noise. The filter operation amounts to replacing the value for each pixel with a weighted average of the values of that pixel and the neighboring pixels in a predefined neighborhood or window (e.g., a 9 9 window). The filter weights depend partly on pixel values and partly on the window size. The present FPGA implementation of a bilateral subtraction filter utilizes a 9 9 window. This implementation was designed to take advantage of the ability to do many of the component computations in parallel pipelines to enable processing of image data at the rate at which they are generated. The filter can be considered to be divided into the following parts (see figure): a) An image pixel pipeline with a 9 9- pixel window generator, b) An array of processing elements; c) An adder tree; d) A smoothing-and-delaying unit; and e) A subtraction unit. After each 9 9 window is created, the affected pixel data are fed to the processing elements. Each processing element is fed the pixel value for its position in the window as well as the pixel value for the central pixel of the window. The absolute difference between these two pixel values is calculated and used as an address in a lookup table. Each processing element has a lookup table, unique for its position in the window, containing the weight coefficients for the Gaussian function for that position. The pixel value is multiplied by the weight, and the outputs of the processing element are the weight and pixel-value weight product. The products and weights are fed to the adder tree. The sum of the products and the sum of the weights are fed to the divider, which computes the sum of products the sum of weights. The output of the divider is denoted the bilateral smoothed image. The smoothing function is a simple weighted average computed over a 3 3 subwindow centered in the 9 9 window. After smoothing, the image is delayed by an additional amount of time needed to match the processing time for computing the bilateral smoothed image. The bilateral smoothed image is then subtracted from the 3 3 smoothed image to produce the final output. The prototype filter as implemented in a commercially available FPGA processes one pixel per clock cycle. Operation at a clock speed of 66 MHz has been demonstrated, and results of a static timing analysis have been interpreted as suggesting that the clock speed could be increased to as much as 100 MHz.

Huertas, Andres↗

A Mixed Integer Efficient Global Optimization Algorithm with Multiple Infill Strategy - Applied to a Wing Topology Optimization Problem

With the advancement in high performance computing and numerical optimization techniques,engineering design optimization problems are becoming more complex, larger scale,higher fidelity, and computationally more demanding, requiring longer run times than ever before. There exists methodologies and techniques that can address some of these challenges but very few can address all, and most are limited in the extent that these concerns can be addressed. With the goal of addressing such challenging engineering problems, we developed anew optimization framework, named AMIEGO, that combines concepts from surrogate-based optimization approaches, gradient-based numerical methods, Partial Least Squares, evolutionary algorithms, and Branch-and-Bound, providing newer capabilities that were not previouslyperceived. However, the original version of this framework, in the process of adaptive samplingto explore and exploit the design space, finds only a single sample point per iteration. The efforthere builds upon this previously developed optimization framework to include multiple infillsampling capability that combines the concept of generalized expected improvement function,unsupervised learning, and multi-objective evolutionary technique. To demonstrate, AMIEGOwith the multiple infill capability (called AMIEGO-MIMOS) solves a series of increasingly difficultengineering design optimization problems. The results reveal the performance of the newapproach is problem dependent. When applied to a ten-bar truss problem, the newly proposedmultiple infill strategy consistently leads to a better design solutions when compared to theexisting CPTV method (implemented with the context of the AMIEGO framework). On theother hand, when applied to a mixed-integer high fidelity wing topology optimization problem- MIMOS, despite showing a steeper convergence at the start, eventually leads to an inferiorsolution as compared to CPTV approach. These results also reveal that a small number ofstarting points, in general, are sufficient to lead to a good overall solution.

Mixed-integer optimization↗

A parallel finite-difference method for computational aerodynamics

A finite-difference scheme for solving complex three-dimensional aerodynamic flow on parallel-processing supercomputers is presented. The method consists of a basic flow solver with multigrid convergence acceleration, embedded grid refinements, and a zonal equation scheme. Multitasking and vectorization have been incorporated into the algorithm. Results obtained include multiprocessed flow simulations from the Cray X-MP and Cray-2. Speedups as high as 3.3 for the two-dimensional case and 3.5 for segments of the three-dimensional case have been achieved on the Cray-2. The entire solver attained a factor of 2.7 improvement over its unitasked version on the Cray-2. The performance of the parallel algorithm on each machine is analyzed.

Swisshelm, Julie M.↗

Mirroring within the Fokker-Planck formulation of cosmic ray pitch angle scattering in homogeneous magnetic turbulence

The Fokker-Planck coefficient for pitch angle scattering, appropriate for cosmic rays in homogeneous, stationary, magnetic turbulence, is computed from first principles. No assumptions are made concerning any special statistical symmetries the random field may have. This result can be used to compute the parallel diffusion coefficient for high energy cosmic rays moving in strong turbulence, or low energy cosmic rays moving in weak turbulence. Becuase of the generality of the magnetic turbulence which is allowed in this calculation, special interplanetary magnetic field features such as discontinuities, or particular wave modes, can be included rigorously. The reduction of this results to previously available expressions for the pitch angle scattering coefficient in random field models with special symmetries is discussed. The general existance of a Dirac delta function in the pitch angle scattering coefficient is demonstrated. It is proved that this delta function is the Fokker-Planck prediction for pitch angle scattering due to mirroring in the magnetic field.

Goldstein, M. L.↗

A 30-day forecast experiment with the GISS model and updated sea surface temperatures

The GISS model was used to compute two parallel global 30-day forecasts for the month January 1974. In one forecast, climatological January sea surface temperatures were used, while in the other observed sea temperatures were inserted and updated daily. A comparison of the two forecasts indicated no clear-cut beneficial effect of daily updating of sea surface temperatures. Despite the rapid decay of daily predictability, the model produced a 30-day mean forecast for January 1974 that was generally superior to persistence and climatology when evaluated over either the globe or the Northern Hemisphere, but not over smaller regions.

Spar, J.↗

Mirroring in the Fokker-Planck coefficient for cosmic-ray pitch-angle scattering in homogeneous magnetic turbulence

The Fokker-Planck coefficient for pitch-angle scattering, appropriate for cosmic rays in homogeneous stationary magnetic turbulence is computed without making any specific assumptions concerning the statistical symmetries of the random field. The Fokker-Planck coefficient obtained can be used to compute the parallel diffusion coefficient for high-energy cosmic rays propagating in the presence of strong turbulence, or for low-energy cosmic rays in the presence of weak turbulence. Because of the generality of magnetic turbulence allowed for in the analysis, special interplanetary magnetic field features, such as discontinuities or particular wave modes, can be included rigorously.

Goldstein, M. L.↗

Basic cluster compression algorithm

Feature extraction and data compression of LANDSAT data is accomplished by BCCA program which reduces costs associated with transmitting, storing, distributing, and interpreting multispectral image data. Algorithm uses spatially local clustering to extract features from image data to describe spectral characteristics of data set. Approach requires only simple repetitive computations, and parallel processing can be used for very high data rates. Program is written in FORTRAN IV for batch execution and has been implemented on SEL 32/55.

Hilbert, E. E.↗

Application of the multigrid method to grid generation

The multigrid method (MGM), used to numerically solve the pair of nonlinear elliptic equations commonly used to generate two dimensional boundary-fitted coordinate systems is discussed. Two different geometries are considered: one involving a coordinate system fitted about a circle and the other selected for an impinging jet flow problem. Two different relaxation schemes are tried: one is successive point overrelaxation and the other is a four-color scheme vectorizeable to take advantage of a parallel processor computer for greater computational speed. Results using MGM are compared with those using SOR (doing successive overrelaxations with the corresponding relaxation scheme on the fine grid only). It is found that MGM becomes significantly more effective than SOR as more accuracy is demanded and as more corrective grids, or more grid points, are used. For the accuracy required, it is found that MGM is two to three times faster than SOR in computing time. With the four-color relaxation scheme as applied to the impinging jet problem, the advantage of MGM over SOR is not as great. This may be due to the effect of a poor initial guess on MGM for this problem.

Ohring, S.↗

A linear decomposition method for large optimization problems. Blueprint for development

A method is proposed for decomposing large optimization problems encountered in the design of engineering systems such as an aircraft into a number of smaller subproblems. The decomposition is achieved by organizing the problem and the subordinated subproblems in a tree hierarchy and optimizing each subsystem separately. Coupling of the subproblems is accounted for by subsequent optimization of the entire system based on sensitivities of the suboptimization problem solutions at each level of the tree to variables of the next higher level. A formalization of the procedure suitable for computer implementation is developed and the state of readiness of the implementation building blocks is reviewed showing that the ingredients for the development are on the shelf. The decomposition method is also shown to be compatible with the natural human organization of the design process of engineering systems. The method is also examined with respect to the trends in computer hardware and software progress to point out that its efficiency can be amplified by network computing using parallel processors.

Sobieszczanski-Sobieski, J.↗

Aerothermal loads analysis for high speed flow over a quilted surface configuration

Attention is given to hypersonic laminar flow over a quilted surface configuration that simulates an array of Space Shuttle Thermal Protection System panels bowed in a spherical shape as a result of thermal gradient through the panel thickness. Pressure and heating loads to the surface are determined. The flow field over the configuration was mathematically modeled by means of time-dependent, three-dimensional conservation of mass, momentum, and energy equations. A boundary mapping technique was then used to obtain a rectangular, parallel piped computational domain, and an explicit MacCormack (1972) explicit time-split predictor corrector finite difference algorithm was used to obtain steady state solutions. Total integrated heating loads vary linearly with bowed height when this value does not exceed the local boundary layer thickness.

Olsen, G. C.↗

Parallel structures in human and computer memory

If one thinks of our experiences as being recorded continuously on film, then human memory can be compared to a film library that is indexed by the contents of the film strips stored in it. Moreover, approximate retrieval cues suffice to retrieve information stored in this library. One recognizes a familiar person in a fuzzy photograph or a familiar tune played on a strange instrument. A computer memory that would allow a computer to recognize patterns and to recall sequences the way humans do is constructed. Such a memory is remarkably similiar in structure to a conventional computer memory and also to the neural circuits in the cortex of the cerebellum of the human brain. It is concluded that the frame problem of artificial intelligence could be solved by the use of such a memory if one were able to encode information about the world properly.

Kanerva, P.↗