Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Iterative methods for large scale static analysis of structures on a scalable multiprocessor supercomputer

A parallel Preconditioned Conjugate Gradient (PCG) iterative solver has been developed and implemented on the iPSC-860 scalable hypercube. This new implementation makes use of the Parallel Automated Runtime Toolkit at ICASE (PARTI) primitives to efficiently program irregular communications patterns that exist in general sparse matrices and in particular in the finite element sparse stiffness matrices. The iterative PCG has been used to solve the finite element equations that result from discretizing large scale aerospace structures. In particular, the static response of the High Speed Civil Transport (HSCT) finite element model is solved on the iPSC-860.

Sobh, Nahil Atef↗

Integrated numerical methods for hypersonic aircraft cooling systems analysis

Numerical methods have been developed for the analysis of hypersonic aircraft cooling systems. A general purpose finite difference thermal analysis code is used to determine areas which must be cooled. Complex cooling networks of series and parallel flow can be analyzed using a finite difference computer program. Both internal fluid flow and heat transfer are analyzed, because increased heat flow causes a decrease in the flow of the coolant. The steady state solution is a successive point iterative method. The transient analysis uses implicit forward-backward differencing. Several examples of the use of the program in studies of hypersonic aircraft and rockets are provided.

Petley, Dennis H.↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment-distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

New computing systems, future computing environment, and their implications on structural analysis and design

Recent advances in computer technology that are likely to impact structural analysis and design of flight vehicles are reviewed. A brief summary is given of the advances in microelectronics, networking technologies, and in the user-interface hardware and software. The major features of new and projected computing systems, including high performance computers, parallel processing machines, and small systems, are described. Advances in programming environments, numerical algorithms, and computational strategies for new computing systems are reviewed. The impact of the advances in computer technology on structural analysis and the design of flight vehicles is described. A scenario for future computing paradigms is presented, and the near-term needs in the computational structures area are outlined.

Noor, Ahmed K.↗

High speed civil transport aerodynamic optimization

This is a report of work in support of the Computational Aerosciences (CAS) element of the Federal HPCC program. Specifically, CFD and aerodynamic optimization are being performed on parallel computers. The long-range goal of this work is to facilitate teraflops-rate multidisciplinary optimization of aerospace vehicles. This year's work is targeted for application to the High Speed Civil Transport (HSCT), one of four CAS grand challenges identified in the HPCC FY 1995 Blue Book. This vehicle is to be a passenger aircraft, with the promise of cutting overseas flight time by more than half. To meet fuel economy, operational costs, environmental impact, noise production, and range requirements, improved design tools are required, and these tools must eventually integrate optimization, external aerodynamics, propulsion, structures, heat transfer, controls, and perhaps other disciplines. The fundamental goal of this project is to contribute to improved design tools for U.S. industry, and thus to the nation's economic competitiveness.

Ryan, James S.↗

Algorithms for parallel flow solvers on message passing architectures

The purpose of this project has been to identify and test suitable technologies for implementation of fluid flow solvers -- possibly coupled with structures and heat equation solvers -- on MIMD parallel computers. In the course of this investigation much attention has been paid to efficient domain decomposition strategies for ADI-type algorithms. Multi-partitioning derives its efficiency from the assignment of several blocks of grid points to each processor in the parallel computer. A coarse-grain parallelism is obtained, and a near-perfect load balance results. In uni-partitioning every processor receives responsibility for exactly one block of grid points instead of several. This necessitates fine-grain pipelined program execution in order to obtain a reasonable load balance. Although fine-grain parallelism is less desirable on many systems, especially high-latency networks of workstations, uni-partition methods are still in wide use in production codes for flow problems. Consequently, it remains important to achieve good efficiency with this technique that has essentially been superseded by multi-partitioning for parallel ADI-type algorithms. Another reason for the concentration on improving the performance of pipeline methods is their applicability in other types of flow solver kernels with stronger implied data dependence. Analytical expressions can be derived for the size of the dynamic load imbalance incurred in traditional pipelines. From these it can be determined what is the optimal first-processor retardation that leads to the shortest total completion time for the pipeline process. Theoretical predictions of pipeline performance with and without optimization match experimental observations on the iPSC/860 very well. Analysis of pipeline performance also highlights the effect of uncareful grid partitioning in flow solvers that employ pipeline algorithms. If grid blocks at boundaries are not at least as large in the wall-normal direction as those immediately adjacent to them, then the first processor in the pipeline will receive a computational load that is less than that of subsequent processors, magnifying the pipeline slowdown effect. Extra compensation is needed for grid boundary effects, even if all grid blocks are equally sized.

Vanderwijngaart, Rob F.↗

Multithreaded Model for Dynamic Load Balancing Parallel Adaptive PDE Computations

We present a multithreaded model for the dynamic load-balancing of numerical, adaptive computations required for the solution of Partial Differential Equations (PDE's) on multiprocessors. Multithreading is used as a means of exploring concurrency in the processor level in order to tolerate synchronization costs inherent to traditional (non-threaded) parallel adaptive PDE solvers. Our preliminary analysis for parallel, adaptive PDE solvers indicates that multithreading can be used an a mechanism to mask overheads required for the dynamic balancing of processor workloads with computations required for the actual numerical solution of the PDE's. Also, multithreading can simplify the implementation of dynamic load-balancing algorithms, a task that is very difficult for traditional data parallel adaptive PDE computations. Unfortunately, multithreading does not always simplify program complexity, often makes code re-usability not an easy task, and increases software complexity.

Chrisochoides, Nikos↗

Space Station Freedom Central Thermal Control System Evolution

The objective of the evolution study is to review the proposed growth scenarios for Space Station Freedom and identify the major CTCS hardware scars and software hooks required to facilitate planned growth and technology obsolescence. The Station's two leading evolutionary configurations are: (1) the Research and Development node, where the fundamental mission is scientific research and commercial endeavors, and (2) the Transportation node, where the emphasis is on supporting Lunar and Mars human exploration. These two nodes evolve from the from the assembly complete configuration by the addition of manned modules, pocket labs, resource nodes, attached payloads, customer servicing facility, and an upper and lower keel and boom truss structure. In the case of the R & D node, the role of the dual keel will be to support external payloads for scientific research. In the case of the Transportation node, the keel will support the Lunar (LTV) and Mars (MTV) transportation vehicle service facilities In addition to external payloads. The transverse boom is extended outboard of the alpha gimbal to accommodate the new solar dynamic arrays for power generation, which will supplement the photovoltaic system. The design, development, deployment, and operation of SSF will take place over a 30 year time period and new Innovations and maturation in technologies can be expected. Evolutionary planning must include the obsolescence and insertion of the new technologies over the life of the program, and the technology growth issues must be addressed in parallel with the development of the baseline thermal control system. Technologies that mature and are available within the next 10 years are best suited for evolutionary consideration as the growth phase begins in the year 2000. To increase TCS capability to accommodate growth using baseline technology would require some penalty in mass, volume, EVA time, manifesting, and operational support. To be cost effective the capabilities of the heat acquisition, transport, and rejection subsystems must be increased.

Bullock, Richard↗

Collaborative Aerospace Research and Fellowship Program at NASA Glenn Research Center

During the summer of 2004, a 10-week activity for university faculty entitled the NASA-OAI Collaborative Aerospace Research and Fellowship Program (CFP) was conducted at the NASA Glenn Research Center in collaboration with the Ohio Aerospace Institute (OAI). This is a companion program to the highly successful NASA Faculty Fellowship Program and its predecessor, the NASA-ASEE Summer Faculty Fellowship Program that operated for 38 years at Glenn. The objectives of CFP parallel those of its companion, viz., (1) to further the professional knowledge of qualified engineering and science faculty,(2) to stimulate an exchange of ideas between teaching participants and employees of NASA, (3) to enrich and refresh the research and teaching activities of participants institutions, and (4) to contribute to the research objectives of Glenn. However, CFP, unlike the NASA program, permits faculty to be in residence for more than two summers and does not limit participation to United States citizens. Selected fellows spend 10 weeks at Glenn working on research problems in collaboration with NASA colleagues and participating in related activities of the NASA-ASEE program. This year's program began officially on June 1, 2004 and continued through August 7, 2004. Several fellows had program dates that differed from the official dates because university schedules vary and because some of the summer research projects warranted a time extension beyond the 10 weeks for satisfactory completion of the work. The stipend paid to the fellows was $1200 per week and a relocation allowance of $1000 was paid to those living outside a 50-mile radius of the Center. In post-program surveys from this and previous years, the faculty cited numerous instances where participation in the program has led to new courses, new research projects, new laboratory experiments, and grants from NASA to continue the work initiated during the summer. Many of the fellows mentioned amplifying material, both in undergraduate and graduate courses, on the basis of the summer s experience at Glenn. A number of 2004 fellows indicated that proposals to NASA will grow out of their summer research projects. In addition, some journal articles and NASA publications will result from this past summer s activities. Fellows from past summers continue to send reprints of articles that resulted from work initiated at Glenn. This report is intended primarily to summarize the research activities comprising the 2004 CFP Program at Glenn. Particular research studies include: 1) Development of an Imaging-Based, Computational Fluid Dynamics Tool to Assess Fluid Mechanics in Experimental Models that Simulate Blood Vessels; 2) Analysis of Nanomaterials Produced from Precursors; and 3) LEO Propagation Analysis Tool.

Heyward, Ann O.↗

Satellite Image Mosaic Engine

A computer program automatically builds large, full-resolution mosaics of multispectral images of Earth landmasses from images acquired by Landsat 7, complete with matching of colors and blending between adjacent scenes. While the code has been used extensively for Landsat, it could also be used for other data sources. A single mosaic of as many as 8,000 scenes, represented by more than 5 terabytes of data and the largest set produced in this work, demonstrated what the code could do to provide global coverage. The program first statistically analyzes input images to determine areas of coverage and data-value distributions. It then transforms the input images from their original universal transverse Mercator coordinates to other geographical coordinates, with scaling. It applies a first-order polynomial brightness correction to each band in each scene. It uses a data-mask image for selecting data and blending of input scenes. Under control by a user, the program can be made to operate on small parts of the output image space, with check-point and restart capabilities. The program runs on SGI IRIX computers. It is capable of parallel processing using shared-memory code, large memories, and tens of central processing units. It can retrieve input data and store output data at locations remote from the processors on which it is executed.

Plesea, Lucian↗

An Experimental Study of Upward Burning Over Long Solid Fuels: Facility Development and Comparison

As NASA's mission evolves, new spacecraft and habitat environments necessitate expanded study of materials flammability. Most of the upward burning tests to date, including the NASA standard material screening method NASA-STD-6001, have been conducted in small chambers where the flame often terminates before a steady state flame is established. In real environments, the same limitations may not be present. The use of long fuel samples would allow the flames to proceed in an unhindered manner. In order to explore sample size and chamber size effects, two large chambers were developed at NASA GRC under the Flame Prevention, Detection and Suppression (FPDS) project. The first was an existing vacuum facility, VF-13, located at NASA John Glenn Research Center. This 6350 liter chamber could accommodate fuels sample lengths up to 2 m. However, operational costs and restricted accessibility limited the test program, so a second laboratory scale facility was developed in parallel. By stacking additional two chambers on top of an existing combustion chamber facility, this 81 liter Stacked-chamber facility could accommodate a 1.5 m sample length. The larger volume, more ideal environment of VF-13 was used to obtain baseline data for comparison with the stacked chamber facility. In this way, the stacked chamber facility was intended for long term testing, with VF-13 as the proving ground. Four different solid fuels (adding machine paper, poster paper, PMMA plates, and Nomex fabric) were tested with fuel sample lengths up to 2 m. For thin samples (papers) with widths up to 5 cm, the flame reached a steady state length, which demonstrates that flame length may be stabilized even when the edge effects are reduced. For the thick PMMA plates, flames reached lengths up to 70 cm but were highly energetic and restricted by oxygen depletion. Tests with the Nomex fabric confirmed that the cyclic flame phenomena, observed in small facility tests, continued over longer sample. New features were also observed at the higher oxygen/pressure conditions available in the large chamber. Comparison of flame behavior between the two facilities under identical conditions revealed disparities, both qualitative and quantitative. This suggests that, in certain ranges of controlling parameters, chamber size and shape could be one of the parameters that affect the material flammability. If this proves to be true, it may limit the applicability of existing flammability data.

Kleinhenz, Julie↗

Can Egypt Become Self-Sufficient in Wheat?

Egypt produces half of the 20 million tons of wheat that it consumes with irrigation and imports the other half. Egypt is also the world's largest importer of wheat. The population of Egypt is currently growing at 2.2% annually, and projections indicate that the demand for wheat will triple by the end of the century. Combining multi-crop and -climate models for different climate change scenarios with recent trends in technology, we estimated that future wheat yield will decline mostly from climate change, despite some yield improvements from new technologies. The growth stimulus from elevated atmospheric CO2 will be overtaken by the negative impact of rising temperatures on crop growth and yield. An ongoing program to double the irrigated land area by 2035 in parallel with crop intensification could increase wheat production and make Egypt self-sufficient in the near future, but would be insufficient after 2040s, even with modest population growth. Additionally, the demand for irrigation will increase from 6 to 20 billion m3 for the expanded wheat production, but even more water is needed to account for irrigation efficiency and salt leaching (to a total of up to 29 billion m3). Supplying water for future irrigation and producing sufficient grain will remain challenges for Egypt.

Asseng, Senthold↗

Charon Toolkit for Parallel, Implicit Structured-Grid Computations: Functional Design

Charon is a software toolkit that enables engineers to develop high-performing message-passing programs in a convenient and piecemeal fashion. Emphasis is on rapid program development and prototyping. In this report a detailed description of the functional design of the toolkit is presented. It is illustrated by the stepwise parallelization of two representative code examples.

VanderWijngaart, Rob F.↗

Performance Comparison of HPF and MPI Based NAS Parallel Benchmarks

Compilers supporting High Performance Form (HPF) features first appeared in late 1994 and early 1995 from Applied Parallel Research (APR), Digital Equipment Corporation, and The Portland Group (PGI). IBM introduced an HPF compiler for the IBM RS/6000 SP2 in April of 1996. Over the past two years, these implementations have shown steady improvement in terms of both features and performance. The performance of various hardware/ programming model (HPF and MPI) combinations will be compared, based on latest NAS Parallel Benchmark results, thus providing a cross-machine and cross-model comparison. Specifically, HPF based NPB results will be compared with MPI based NPB results to provide perspective on performance currently obtainable using HPF versus MPI or versus hand-tuned implementations such as those supplied by the hardware vendors. In addition, we would also present NPB, (Version 1.0) performance results for the following systems: DEC Alpha Server 8400 5/440, Fujitsu CAPP Series (VX, VPP300, and VPP700), HP/Convex Exemplar SPP2000, IBM RS/6000 SP P2SC node (120 MHz), NEC SX-4/32, SGI/CRAY T3E, and SGI Origin2000. We would also present sustained performance per dollar for Class B LU, SP and BT benchmarks.

Saini, Subhash↗

Parallel processors and nonlinear structural dynamics algorithms and software

A nonlinear structural dynamics finite element program was developed to run on a shared memory multiprocessor with pipeline processors. The program, WHAMS, was used as a framework for this work. The program employs explicit time integration and has the capability to handle both the nonlinear material behavior and large displacement response of 3-D structures. The elasto-plastic material model uses an isotropic strain hardening law which is input as a piecewise linear function. Geometric nonlinearities are handled by a corotational formulation in which a coordinate system is embedded at the integration point of each element. Currently, the program has an element library consisting of a beam element based on Euler-Bernoulli theory and trianglar and quadrilateral plate element based on Mindlin theory.

Belytschko, Ted↗

Massively parallel processor

A brief description is given of the Massively Parallel Processor (MPP). Major applications of the MPP are in the area of image processing (where the operands are often very small integers) from very high spatial resolution passive image sensors, signal processing of radar data, and numerical modeling simulations of climate. The system can be programmed in assembly language or a high level language. Information on background, status, architecture, programming, hardware reliability, applications, and the MPP's development as a national resource for parallel algorithm research are presented in outline form.

Source record↗

Multiprogramming performance degradation - Case study on a shared memory multiprocessor

The performance degradation due to multiprogramming overhead is quantified for a parallel-processing machine. Measurements of real workloads were taken, and it was found that there is a moderate correlation between the completion time of a program and the amount of system overhead measured during program execution. Experiments in controlled environments were then conducted to calculate a lower bound on the performance degradation of parallel jobs caused by multiprogramming overhead. The results show that the multiprogramming overhead of parallel jobs consumes at least 4 percent of the processor time. When two or more serial jobs are introduced into the system, this amount increases to 5.3 percent

Dimpsey, R. T.↗