Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Multi-Zone Liquid Thrust Chamber Performance Code with Domain Decomposition for Parallel Processing

Computational Fluid Dynamics (CFD) has considerably evolved in the last decade. There are many computer programs that can perform computations on viscous internal or external flows with chemical reactions. CFD has become a commonly used tool in the design and analysis of gas turbines, ramjet combustors, turbo-machinery, inlet ducts, rocket engines, jet interaction, missile, and ramjet nozzles. One of the problems of interest to NASA has always been the performance prediction for rocket and air-breathing engines. Due to the complexity of flow in these engines it is necessary to resolve the flowfield into a fine mesh to capture quantities like turbulence and heat transfer. However, calculation on a high-resolution grid is associated with a prohibitively increasing computational time that can downgrade the value of the CFD for practical engineering calculations. The Liquid Thrust Chamber Performance (LTCP) code was developed for NASA/MSFC (Marshall Space Flight Center) to perform liquid rocket engine performance calculations. This code is a 2D/axisymmetric full Navier-Stokes (NS) solver with fully coupled finite rate chemistry and Eulerian treatment of liquid fuel and/or oxidizer droplets. One of the advantages of this code has been the resemblance of its input file to the JANNAF (Joint Army Navy NASA Air Force Interagency Propulsion Committee) standard TDK code, and its automatic grid generation for JANNAF defined combustion chamber wall geometry. These options minimize the learning effort for TDK users, and make the code a good candidate for performing engineering calculations. Although the LTCP code was developed for liquid rocket engines, it is a general-purpose code and has been used for solving many engineering problems. However, the single zone formulation of the LTCP has limited the code to be applicable to problems with complex geometry. Furthermore, the computational time becomes prohibitively large for high-resolution problems with chemistry, two-equation turbulence model, and two-phase flow. To overcome these limitations, the LTCP code is rewritten to include the multi-zone capability with domain decomposition that makes it suitable for parallel processing, i.e., enabling the code to run every zone or sub-domain on a separate processor. This can reduce the run time by a factor of 6 to 8, depending on the problem.

Homayun K. Navaz↗

Debugging tasked Ada programs

The applications for which Ada was developed require distributed implementations of the language and extensive use of tasking facilities. Debugging and testing technology as it applies to parallel features of languages currently falls short of needs. Thus, the development of embedded systems using Ada pose special challenges to the software engineer. Techniques for distributing Ada programs, support for simulating distributed target machines, testing facilities for tasked programs, and debugging support applicable to simulated and to real targets all need to be addressed. A technique is presented for debugging Ada programs that use tasking and it describes a debugger, called AdaTAD, to support the technique. The debugging technique is presented together with the use interface to AdaTAD. The component of AdaTAD that monitors and controls communication among tasks was designed in Ada and is presented through an example with a simple tasked program.

Fainter, R. G.↗

Compile-time estimation of communication costs in multicomputers

An important problem facing numerous research projects on parallelizing compilers for distributed memory machines is that of automatically determining a suitable data partitioning scheme for a program. Any strategy for automatic data partitioning needs a mechanism for estimating the performance of a program under a given partitioning scheme, the most crucial part of which involves determining the communication costs incurred by the program. A methodology is described for estimating the communication costs at compile-time as functions of the numbers of processors over which various arrays are distributed. A strategy is described along with its theoretical basis, for making program transformations that expose opportunities for combining of messages, leading to considerable savings in the communication costs. For certain loops with regular dependences, the compiler can detect the possibility of pipelining, and thus estimate communication costs more accurately than it could otherwise. These results are of great significance to any parallelization system supporting numeric applications on multicomputers. In particular, they lay down a framework for effective synthesis of communication on multicomputers from sequential program references.

Gupta, Manish↗

Performance of a parallel code for the Euler equations on hypercube computers

The performance of hypercubes were evaluated on a computational fluid dynamics problem and the parallel environment issues were considered that must be addressed, such as algorithm changes, implementation choices, programming effort, and programming environment. The evaluation focuses on a widely used fluid dynamics code, FLO52, which solves the two dimensional steady Euler equations describing flow around the airfoil. The code development experience is described, including interacting with the operating system, utilizing the message-passing communication system, and code modifications necessary to increase parallel efficiency. Results from two hypercube parallel computers (a 16-node iPSC/2, and a 512-node NCUBE/ten) are discussed and compared. In addition, a mathematical model of the execution time was developed as a function of several machine and algorithm parameters. This model accurately predicts the actual run times obtained and is used to explore the performance of the code in interesting but yet physically realizable regions of the parameter space. Based on this model, predictions about future hypercubes are made.

Barszcz, Eric↗

A Partitioned - Task Parallel Implementation of the NASA Multiscale Analysis Tool for High Performance Computing

The NASA Multiscale Analysis Tool (NASMAT) is a platform for multiscale modeling of composites which can perform analysis of materials with any arbitrary number of length scales. The platform supports modularity, scalability, and interoperability using recursive procedures and data structures. A Macro solver driven parallelization scheme often limits the capability of NASMAT to scale as it has access to limited memory and number of cores (often one core/thread) and often forces to implement macro solver specific changes to the platform. In this work, a partitioned task-parallel approach is adopted, where the parallelization strategy adopted for NASMAT is independent of the macro solver and the computational resources are managed independently. The programming architecture takes into account the hierarchy of multiple scales (task-dependence) and the heterogeneous nature (dynamic load balancing) of computation through implementation of a hierarchy-informed task parallel model. The partitioned nature of the framework further extends the “plug and play” capability of NASMAT. preCICE, an open-source library for coupling multiphysics solver in a partitioned manner, is adopted to integrate NASMAT with an external macro solver by implementing a NASMAT adapter for preCICE. Speedup and scalability of the framework is studied for micromechanical models of varying size.

task-parallel↗

Facilitating Student Involvement in NASA Research: The NASA Space Grant Aeronautics Example

Many consider NASA programs to be exclusively space-oriented. However, NASA's roots originated in the aeronautical sciences. Recent developments within NASA elevated the declining role of aeronautics back to a position of priority. On a parallel pattern, aeronautics was a priority in the legislation which authorized the National Space Grant College and Fellowship Program. This paper outlines the development of the aeronautics aspect of the National Space Grant College and Fellowship Program, and the resulting student opportunities in research. Results from two aeronautics surveys provide a baseline and direction for further development. A key result of this work is the increase in student research opportunities which now exist in more states and at the national level.

Bowen, Brent D.↗

Energy Usage in an Embedded Space Vision Application on a Tiled Architecture

The need for greater autonomy in platforms such as planetary rovers is driving rapidly to codes that far overwhelm the capabilities of conventional space-qualified single core processors to run them in real-time. However, a new generation of potentially space-qualified 2D "tiled" multi-core microprocessor chips is emerging with significant performance potential. Leveraging such inherently parallel hardware for space platforms requires consideration of both time and power limitations - the latter of which is not normally done in conventional parallel computing. This paper takes one such application, Rockster, and analyzes it for energy usage when ported to a multi-core tiled chip such as may come from the Maestro program. The results demonstrate not only the criticality of memory and interconnect in the energy of real-time parallel codes, but also the effects of possible "energy-aware" changes in partitioning and algorithm design.

multi-core processors↗

Applications of the massively parallel machine, the MasPar MP-1, to Earth sciences

The computational workload of upcoming NASA science missions, especially the ground data processing for the Earth Observing System, is projected to be quite large (in the 50 to 100 gigaFLOPS range) and corespondingly very expensive to perform using conventional supercomputer systems. High performance, general purpose massively parallel computer systems such as the MasPar MP-1 are being investigated by NASA as a more cost effective alternative. Massively parallel systems are targeted for accelerated development and maturation by NASA's upcoming five-year High Performance Computing and Communications Program. A summary of the broad range of applications currently running on the MP-1 at NASA/Goddard are presented in this paper along with descriptions of the parallel algorithmic techniques employed in five applications that have bearing on Earth sciences.

Fischer, James R.↗

Innovative Language-Based & Object-Oriented Structured AMR Using Fortran 90 and OpenMP

Parallel adaptive mesh refinement (AMR) is an important numerical technique that leads to the efficient solution of many physical and engineering problems. In this paper, we describe how AMR programing can be performed in an object-oreinted way using the modern aspects of Fortran 90 combined with the parallelization features of OpenMP.

software↗

PROGRESS REPORT ON THE DEVELOPMENT OF PROTECTED CONSTRUCTION FOR HYPERSONIC VEHICLES

The structural problems of re-entry occasioned by aerodynamic heating are now generally well known. Typical of these heating problems is the environment experienced by the manned lifting vehicle re-entering from a low altitude orbit. Figure 1 shows typical curves of lower surface temperature as a function of time for a wing loading of 25 Ibs/ft(exp 2), and a lift-drag ratio of 2.5. The two curves cover a practical range of re-entry angles, the lower curve representing an ideal re-entry path with zero dive angle and the upper curve assuming a 2° error in the re-entry angle. Maneuvers for course change or correction are also added to the upper curve. Significant factors from this curve are the relationship between the maximum temperatures and the capabilities of available metallic material, and also the long re-entry time and its effect on total heat load. If a very shallow re-entry is made, maximum lower surface temperatures reach 2000°F which is just within the range of conventional superalloys. To accommodate practical re-entry angles and maneuvers, however, the temperature capability must reach about 2500°F which requires refractory metals. The re-entry time may be as high as 100 minutes, which gives total heat loads of approximately 40,000 BTU/ft(exp 2). A heat load of this magnitude, with equilibrium surface temperatures of the values shown, suggests that a lighter airframe can be constructed by dissipating the heat by radiation from the surface, rather than by absorbing it with a heat sink, or surface cooling, or ablation. Air Force programs to provide airframes for this type of environment have involved the parallel development of a number of different structural concepts. The development to be discussed here was carried out by Bell Aerosystems Company for the Fabrication Branch, Manufacturing Technology Laboratory, Directorate of Materials and Processes, Aeronautical Systems' Division, Wright-Patterson Air Force Base, Ohio, under Contract AF33(600)-40100 (Double-Wall Cooled Structure).

Reentry vehicle↗

Thrust imbalance of the Space Shuttle solid rocket motors

The Monte Carlo statistical analysis of thrust imbalance is applied to both the Titan IIIC and the Space Shuttle solid rocket motors (SRMs) firing in parallel, and results are compared with those obtained from the Space Shuttle program. The test results are examined in three phases: (1) pairs of SRMs selected from static tests of the four developmental motors (DMs 1 through 4); (2) pairs of SRMs selected from static tests of the three quality assurance motors (QMs 1 through 3); (3) SRMs on the first flight test vehicle (STS-1A and STS-1B). The simplified internal ballistic model utilized for computing thrust from head-end pressure measurements on flight tests is shown to agree closely with measured thrust data. Inaccuracies in thrust imbalance evaluation are explained by possible flight test instrumentation errors.

Foster, W. A., Jr.↗

Application of the p-version of the finite-element method to global-local problems

A brief survey is given of some recent developments in finite-element analysis technology which bear upon the three main research areas under consideration in this workshop: (1) analysis methods; (2) software testing and quality assurance; and (3) parallel processing. The variational principle incorporated in a finite-element computer program, together with a particular set of input data, determines the exact solution corresponding to that input data. Most finite-element analysis computer programs are based on the principle of virtual work. In the following, researchers consider only programs based on the principle of virtual work and denote the exact displacement vector field corresponding to some specific set of input data by vector u(EX). The exact solution vector u(EX) is independent of the design of the mesh or the choice of elements. Except for very simple problems, or specially constructed test problems, vector u(EX) is not known. Researchers perform a finite-element analysis (or any other numerical analysis) because they wish to make conclusions concerning the response of a physical system to certain imposed conditions, as if vector u(EX) were known.

Szabo, Barna A.↗

The PARTY parallel runtime system

In the present automated system for the organization of the data and computational operations entailed by parallel problems, in ways that optimize multiprocessor performance, general heuristics for partitioning program data and control are implemented by capturing and manipulating representations of a computation at run time. These heuristics are directed toward the dynamic identification and allocation of concurrent work in computations with irregular computational patterns. An optimized static-workload partitioning is computed for such repetitive-computation pattern problems as the iterative ones employed in scientific computation.

Saltz, J. H.↗

Load Variation Influences on Joint Work During Squat Exercise in Reduced Gravity

Resistance exercises that load the axial skeleton, such as the parallel squat, are incorporated as a critical component of a space exercise program designed to maximize the stimuli for bone remodeling and muscle loading. Astronauts on the International Space Station perform regular resistance exercise using the Advanced Resistive Exercise Device (ARED). Squat exercises on Earth entail moving a portion of the body weight plus the added bar load, whereas in microgravity the body weight is 0, so all load must be applied via the bar. Crewmembers exercising in microgravity currently add approx.70% of their body weight to the bar load as compensation for the absence of the body weight. This level of body weight replacement (BWR) was determined by crewmember feedback and personal experience without any quantitative data. The purpose of this evaluation was to utilize computational simulation to determine the appropriate level of BWR in microgravity necessary to replicate lower extremity joint work during squat exercise in normal gravity based on joint work. We hypothesized that joint work would be positively related to BWR load.

DeWitt, John K.↗

Real-Time Cognitive Computing Architecture for Data Fusion in a Dynamic Environment

A novel cognitive computing architecture is conceptualized for processing multiple channels of multi-modal sensory data streams simultaneously, and fusing the information in real time to generate intelligent reaction sequences. This unique architecture is capable of assimilating parallel data streams that could be analog, digital, synchronous/asynchronous, and could be programmed to act as a knowledge synthesizer and/or an "intelligent perception" processor. In this architecture, the bio-inspired models of visual pathway and olfactory receptor processing are combined as processing components, to achieve the composite function of "searching for a source of food while avoiding the predator." The architecture is particularly suited for scene analysis from visual data and odorant.

Duong, Tuan A.↗

Design and fabrication of a stereoscopic rear-viewing endoscopic tool (MARVEL)

The use of minimally invasive neurosurgical techniques has experienced a growing interest among patients and surgeons alike due their the numerous advantages over more traditional operations. Current methods employ the use either rigid or articulating endoscopes, both of which are inserted through a small diameter opening in the skull near the intended surgical region. Although these devices aid the surgeon in viewing the operating site, they are severely limited by their inability to provide dimensional awareness to the user. The Multi-Angle Rear-Viewing Endoscopic TooL (MARVEL) has been designed such that the compact form of existing endoscopes is maintained while also providing the user with an enhanced 3-dimensional stereo view. The design of the alpha prototype for MARVEL was completed over the course of three months. In parallel with development, the device's optics were characterized and several software programs were generated. Once assembled, MARVEL will hold many advantages over existing endoscopic tools. It is expected that the device will greatly outperform its traditional counterparts by both increasing safety and decreasing the duration of surgical procedures. After completion, the device will be demonstrated to the project sponsors at the Skull Base Institute where it will undergo full evaluation and a feasibility analysis.

Strongrich, Andrew↗

Design of a real-time wind turbine simulator using a custom parallel architecture

The design of a new parallel-processing digital simulator is described. The new simulator has been developed specifically for analysis of wind energy systems in real time. The new processor has been named: the Wind Energy System Time-domain simulator, version 3 (WEST-3). Like previous WEST versions, WEST-3 performs many computations in parallel. The modules in WEST-3 are pure digital processors, however. These digital processors can be programmed individually and operated in concert to achieve real-time simulation of wind turbine systems. Because of this programmability, WEST-3 is very much more flexible and general than its two predecessors. The design features of WEST-3 are described to show how the system produces high-speed solutions of nonlinear time-domain equations. WEST-3 has two very fast Computational Units (CU's) that use minicomputer technology plus special architectural features that make them many times faster than a microcomputer. These CU's are needed to perform the complex computations associated with the wind turbine rotor system in real time. The parallel architecture of the CU causes several tasks to be done in each cycle, including an IO operation and the combination of a multiply, add, and store. The WEST-3 simulator can be expanded at any time for additional computational power. This is possible because the CU's interfaced to each other and to other portions of the simulation using special serial buses. These buses can be 'patched' together in essentially any configuration (in a manner very similar to the programming methods used in analog computation) to balance the input/ output requirements. CU's can be added in any number to share a given computational load. This flexible bus feature is very different from many other parallel processors which usually have a throughput limit because of rigid bus architecture.

Hoffman, John A.↗

The Necessity of Functional Analysis for Space Exploration Programs

As NASA moves toward expanded commercial spaceflight within its human exploration capability, there is increased emphasis on how to allocate responsibilities between government and commercial organizations to achieve coordinated program objectives. The practice of program-level functional analysis offers an opportunity for improved understanding of collaborative functions among heterogeneous partners. Functional analysis is contrasted with the physical analysis more commonly done at the program level, and is shown to provide theoretical performance, risk, and safety advantages beneficial to a government-commercial partnership. Performance advantages include faster convergence to acceptable system solutions; discovery of superior solutions with higher commonality, greater simplicity and greater parallelism by substituting functional for physical redundancy to achieve robustness and safety goals; and greater organizational cohesion around program objectives. Risk advantages include avoidance of rework by revelation of some kinds of architectural and contractual mismatches before systems are specified, designed, constructed, or integrated; avoidance of cost and schedule growth by more complete and precise specifications of cost and schedule estimates; and higher likelihood of successful integration on the first try. Safety advantages include effective delineation of must-work and must-not-work functions for integrated hazard analysis, the ability to formally demonstrate completeness of safety analyses, and provably correct logic for certification of flight readiness. The key mechanism for realizing these benefits is the development of an inter-functional architecture at the program level, which reveals relationships between top-level system requirements that would otherwise be invisible using only a physical architecture. This paper describes the advantages and pitfalls of functional analysis as a means of coordinating the actions of large heterogeneous organizations for space exploration programs.

program management↗