Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38

LAURA Users Manual: 5.6

This users manual provides in-depth information concerning installation and execution of Laura, version 5. Laura is a structured, multiblock, computational aerothermodynamic simulation code. Version 5 represents a major refactoring of the original Fortran 77 Laura code toward a modular structure afforded by Fortran 95. The refactoring improved usability and maintainability by eliminating the requirement for problem-dependent recompilations, providing more intuitive distribution of functionality, and simplifying inter- faces required for multi-physics coupling. As a result, Laura now shares gas-physics modules, MPI modules, and other low-level modules with the Fun3D unstructured-grid code. In addition to internal refactoring, several new features and capabilities have been added, e.g., a GNU-standard installation process, parallel load balancing, automatic trajectory point sequencing, free-energy minimization, and coupled ablation and flow field radiation.

Aerodynamics↗

Numerical results for the thermal scattering functions

A recent formulation in radiative transfer defined the thermal scattering functions that characterize radiative transfer from a general plane-parallel finite medium driven solely by an internal distribution of thermal sources. Exiting diffuse intensities are expressed as space convolutions of the thermal scattering functions with any thermal source distribution. A parametric study is presented to obtain the basic structure of these scattering functions. The independent variables of these azimuthally independent functions are the direction cosine and source location, while the parameters are the single-scattering albedo, total optical depth, and the asymmetry factor in the Henyey-Greenstein phase function. The basic functional trends are discussed by using various parametric plots, and selected results are given to allow numerical checks. The computational method is invariant imbedding.

Cogley, A. C.↗

Analysis of turbulent free-jet hydrogen-air diffusion flames with finite chemical reaction rates

A numerical analysis is presented of the nonequilibrium flow field resulting from the turbulent mixing and combustion of an axisymmetric hydrogen jet in a supersonic parallel ambient air stream. The effective turbulent transport properties are determined by means of a two-equation model of turbulence. The finite-rate chemistry model considers eight elementary reactions among six chemical species: H, O, H2O, OH, O2 and H2. The governing set of nonlinear partial differential equations was solved by using an implicit finite-difference procedure. Radial distributions were obtained at two downstream locations for some important variables affecting the flow development, such as the turbulent kinetic energy and its dissipation rate. The results show that these variables attain their peak values on the axis of symmetry. The computed distribution of velocity, temperature, and mass fractions of the chemical species gives a complete description of the flow field. The numerical predictions were compared with two sets of experimental data. Good qualitative agreement was obtained.

Sislian, J. P.↗

LAURA Users Manual: 5.3-48528

This users manual provides in-depth information concerning installation and execution of LAURA, version 5. LAURA is a structured, multi-block, computational aerothermodynamic simulation code. Version 5 represents a major refactoring of the original Fortran 77 LAURA code toward a modular structure afforded by Fortran 95. The refactoring improved usability and maintainability by eliminating the requirement for problem-dependent re-compilations, providing more intuitive distribution of functionality, and simplifying interfaces required for multi-physics coupling. As a result, LAURA now shares gas-physics modules, MPI modules, and other low-level modules with the FUN3D unstructured-grid code. In addition to internal refactoring, several new features and capabilities have been added, e.g., a GNU-standard installation process, parallel load balancing, automatic trajectory point sequencing, free-energy minimization, and coupled ablation and flowfield radiation.

Mazaheri, Alireza↗

Computations of Boiling in Microgravity

The absence (or reduction) of gravity, can lead to major changes in boiling heat transfer. On Earth, convection has a major effect on the heat distribution ahead of an evaporation front, and buoyancy determines the motion of the growing bubbles. In microgravity, convection and buoyancy are absent or greatly reduced and the dynamics of the growing vapor bubbles can change in a fundamental way. In particular, the lack of redistribution of heat can lead to a large superheat and explosive growth of bubbles once they form. While considerable efforts have been devoted to examining boiling experimentally, including the effect of microgravity, theoretical and computational work have been limited. Here, the growth of boiling bubbles is studied by direct numerical simulations where the flow field is fully resolved and the effects of inertia, viscosity, surface deformation, heat conduction and convection, as well as the phase change, are fully accounted for. Boiling involves both fluid flow and heat transfer and thus requires the solution of the Navier-Stokes and the energy equations. The numerical method is based on writing one set of governing transport equations which is valid in both the liquid and vapor phases. This local, single-field formulation incorporates the effect of the interface in the governing equations as source terms acting only at the interface. These sources account for surface tension and latent heat in the equations for conservation of momentum and energy as well as mass transfer across the interface due to phase change. The single-field formulation naturally incorporates the correct mass, momentum and energy balances across the interface. Integration of the conservation equations across the interface directly yields the jump conditions derived in the local instant formulation for two-phase systems. In the numerical implementation, the conservation equations for the whole computational domain (both vapor and liquid) are solved using a stationary grid and the phase boundary is followed by a moving unstructured two-dimensional grid. While two-dimensional simulations have been used for preliminary studies and to examine the resolution requirement, the focus is on fully three-dimensional simulations. The numerical methodology, including the parallelization and grid refinement strategy is discussed, and preliminary results shown. For buoyancy driven flow, the heat transfer is in good agreement with experimental correlations. The changes when gravity is turned off and/or fluid shear is added are discussed, as well as the difference between simulations of a layer freely releasing bubbles versus simulations using only one wavelength initial perturbation. Figure 1 shows the early stages of the formation of a three-dimensional bubble from a thin vapor layer. The boundary conditions are periodic in the x and y direction, the bottom is a hot and the top allows a free outflow. The jagged edge of the surface close to the bottom of the computational domain is due to some of the surface elements being on the other side of the domain and some elements not plotted by our plotting routine. In the second figure, we show the temperature distribution through two perpendicular planes.

Tryggvason, G.↗

Solar flare model atmospheres

Solar flare model atmospheres computed under the assumption of energetic equilibrium in the chromosphere are presented. The models use a static, one-dimensional plane parallel geometry and are designed within a physically self-consistent coronal loop. Assumed flare heating mechanisms include collisions from a flux of non-thermal electrons and x-ray heating of the chromosphere by the corona. The heating by energetic electrons accounts explicitly for variations of the ionized fraction with depth in the atmosphere. X-ray heating of the chromosphere by the corona incorporates a flare loop geometry by approximating distant portions of the loop with a series of point sources, while treating the loop leg closest to the chromospheric footpoint in the plane-parallel approximation. Coronal flare heating leads to increased heat conduction, chromospheric evaporation and subsequent changes in coronal pressure; these effects are included self-consistently in the models. Cooling in the chromosphere is computed in detail for the important optically thick HI, CaII and MgII transitions using the non-LTE prescription in the program MULTI. Hydrogen ionization rates from x-ray photo-ionization and collisional ionization by non-thermal electrons are included explicitly in the rate equations. The models are computed in the 'impulsive' and 'equilibrium' limits, and in a set of intermediate 'evolving' states. The impulsive atmospheres have the density distribution frozen in pre-flare configuration, while the equilibrium models assume the entire atmosphere is in hydrostatic and energetic equilibrium. The evolving atmospheres represent intermediate stages where hydrostatic equilibrium has been established in the chromosphere and corona, but the corona is not yet in energetic equilibrium with the flare heating source. Thus, for example, chromospheric evaporation is still in the process of occurring.

Hawley, Suzanne L.↗

Parallel algorithms for interactive manipulation of digital terrain models

Interactive three-dimensional graphics applications, such as terrain data representation and manipulation, require extensive arithmetic processing. Massively parallel machines are attractive for this application since they offer high computational rates, and grid connected architectures provide a natural mapping for grid based terrain models. Presented here are algorithms for data movement on the massive parallel processor (MPP) in support of pan and zoom functions over large data grids. It is an extension of earlier work that demonstrated real-time performance of graphics functions on grids that were equal in size to the physical dimensions of the MPP. When the dimensions of a data grid exceed the processing array size, data is packed in the array memory. Windows of the total data grid are interactively selected for processing. Movement of packed data is needed to distribute items across the array for efficient parallel processing. Execution time for data movement was found to exceed that for arithmetic aspects of graphics functions. Performance figures are given for routines written in MPP Pascal.

Davis, E. W.↗

Location and characteristics of the reconnection X-line deduced from low-altitude satellite and radar observations

We present an analysis of a cusp ion step observed between two poleward-moving events of enhanced ionospheric electron temperature. From the computed variation of the reconnection rate and the onset times of the associated ionospheric events, the distance between the satellite and the X-line can be estimated, but with a large uncertainty due to that in the determination of the low-energy cut-off of the ion velocity distribution function, f(E). Nevertheless, analysis of the time series f(t) shows the reconnection site to be on the dayside magnetopause, consistent with the pulsating cusp model, and the best estimate of the X-line location is 13 R(E) from the satellite. The ion precipitation is used to reconstruct the field-parallel part of the Cowley-D ion distribution function injected into the open low latitude boundary layer (LLBL) in the vicinity of the X-line. From this the Alfven speed, plasma density, magnetic field, parallel ion temperature, and flow velocity of the magnetosheath near the X-line can be derived.

Lockwood, M.↗

Molecular dynamics of the water liquid-vapor interface

The results of molecular dynamics calculations on the equilibrium interface between liquid water and its vapor at 325 K are presented. For the TIP4P model of water intermolecular pair potentials, the average surface dipole density points from the vapor to the liquid. The most common orientations of water molecules have the C2 nu molecular axis roughly parallel to the interface. The distributions are quite broad and therefore compatible with the intermolecular correlations characteristic of bulk liquid water. All near-neighbor pairs in the outermost interfacial layers are hydrogen bonded according to the common definition adopted here. The orientational preferences of water molecules near a free surface differ from those near rigidly planar walls which can be interpreted in terms of patterns found in hexagonal ice 1. The mean electric field in the interfacial region is parallel to the mean polarization which indicates that attention cannot be limited to dipolar charge distributions in macroscopic descriptions of the electrical properties of this interface. The value of the surface tension obtained is 132 +/- 46 dyn/cm, significantly different from the value for experimental water of 68 dyn/cm at 325 K.

NASA Center ARC↗

The PISCES 2 parallel programming environment

PISCES 2 is a programming environment for scientific and engineering computations on MIMD parallel computers. It is currently implemented on a flexible FLEX/32 at NASA Langley, a 20 processor machine with both shared and local memories. The environment provides an extended Fortran for applications programming, a configuration environment for setting up a run on the parallel machine, and a run-time environment for monitoring and controlling program execution. This paper describes the overall design of the system and its implementation on the FLEX/32. Emphasis is placed on several novel aspects of the design: the use of a carefully defined virtual machine, programmer control of the mapping of virtual machine to actual hardware, forces for medium-granularity parallelism, and windows for parallel distribution of data. Some preliminary measurements of storage use are included.

Pratt, Terrence W.↗

Automatic array alignment in data-parallel programs

FORTRAN 90 and other data-parallel languages express parallelism in the form of operations on data aggregates such as arrays. Misalignment of the operands of an array operation can reduce program performance on a distributed-memory parallel machine by requiring nonlocal data accesses. Determining array alignments that reduce communication is therefore a key issue in compiling such languages. We present a framework for the automatic determination of array alignments in array-based, data-parallel languages. Our language model handles array sectioning, reductions, spreads, transpositions, and masked operations. We decompose alignment functions into three constituents: axis, stride, and offset. For each of these subproblems, we show how to solve the alignment problem for a basic block of code, possibly containing common subexpressions. Alignments are generated for all array objects in the code, both named program variables and intermediate results. We assign computation to processors by virtue of explicit alignment of all temporaries; the resulting work assignment is in general better than that provided by the 'owner-computes' rule. Finally, we present some ideas for dealing with control flow, replication, and dynamic alignments that depend on loop induction variables.

Chatterjee, Siddhartha↗

Automatic Data Traffic Control on DSM Architecture

We study data traffic on distributed shared memory machines and conclude that data placement and grouping improve performance of scientific codes. We present several methods which user can employ to improve data traffic in his code. We report on implementation of a tool which detects the code fragments causing data congestions and advises user on improvements of data routing in these fragments. The capabilities of the tool include deduction of data alignment and affinity from the source code; detection of the code constructs having abnormally high cache or TLB misses; generation of data placement constructs. We demonstrate the capabilities of the tool on experiments with NAS parallel benchmarks and with a simple computational fluid dynamics application ARC3D.

Frumkin, Michael↗

Physical and dynamical studies of meteors. Meteor-fragmentation and stream-distribution studies

Population parameters of 275 streams including 20 additional streams in the synoptic-year sample were found by a computer technique. Some 16 percent of the sample is in these streams. Four meteor streams that have close orbital resemblance to Adonis cannot be positively identified as meteors ejected by Adonis within the last 12000 years. Ceplecha's discrete levels of meteor height are not evident in radar meteors. The spread of meteoroid fragments along their common trajectory was computed for most of the observed radar meteors. There is an unexpected relationship between spread and velocity that perhaps conceals relationships between fragmentation and orbits; a theoretical treatment will be necessary to resolve these relationships. Revised unbiased statistics of synoptic-year orbits are presented, together with parallel statistics for the 1961 to 1965 radar meteor orbits.

Sekanina, Z.↗

Parametric analysis of hollow conductor parallel and coaxial transmission lines for high frequency space power distribution

A parametric analysis was performed of transmission cables for transmitting electrical power at high voltage (up to 1000 V) and high frequency (10 to 30 kHz) for high power (100 kW or more) space missions. Large diameter (5 to 30 mm) hollow conductors were considered in closely spaced coaxial configurations and in parallel lines. Formulas were derived to calculate inductance and resistance for these conductors. Curves of cable conductance, mass, inductance, capacitance, resistance, power loss, and temperature were plotted for various conductor diameters, conductor thickness, and alternating current frequencies. An example 5 mm diameter coaxial cable with 0.5 mm conductor thickness was calculated to transmit 100 kW at 1000 Vac, 50 m with a power loss of 1900 W, an inductance of 1.45 micron and a capacitance of 0.07 micron-F. The computer programs written for this analysis are listed in the appendix.

Jeffries, K. S.↗

Solar flare model atmospheres

Solar flare model atmospheres computed under the assumption of energetic equilibrium in the chromosphere are presented. The models use a static, one-dimensional plane-parallel geometry and are designed within a physically self-consistent coronal loop. Assumed flare heating mechanisms include collisions from a flux of nonthermal electrons and X-ray heating of the chromosphere by the corona. The heating by energetic electrons accounts explicitly for variations of the ionized fraction with depth in the atmosphere. X-ray heating of the chromosphere by the corona incorporates a flare loop geometry by approximating distant portions of the loop with a series of point sources, while treating the loop leg closest to the chromospheric footpoint in the plane-parallel approximation. Coronal flare heating leads to increased heat conduction, chromospheric evaporation and subsequent changes in coronal pressure; these effects are included self-consistently in the models. Cooling in the chromosphere is computed in detail for the important optically thick H I, Ca II and Mg II transitions using the non-local thermodynamic equilibrium (non-LTE) prescription in the program MULTI. Hydrogen ionization rates from X-ray photoionization and collisional ionization by nonthermal electrons are included explicitly in the rate equations. The models are computed in the 'impulsive' and 'equilibrium' limits, and in a set of intermediate 'evolving' states. The impulsive atmospheres have the density distribution frozen in the pre-flare configuration, while the equilibrium models assume the entire atmosphere is in hydrostatic and energetic equilibrium. The evolving atmospheres represent intermediate stages where hydrostatic equilibrium has been established in the chromosphere and corona, but the corona is not yet in energetic equilibrium with the flare heating source. Thus, for example, chromospheric evaporation is still in the process of occurring. We have computed the chromospheric radiation that results from a range of coronal heating rates, with particular emphasis on the widely observed diagnostic H(alpha). Our conclusion is that the H(alpha) fluxes and profiles actually observed in flares can only be produced under conditions of a low-pressure corona with strong beam heating. Therefore we suggest that H(alpha) in flares is produced primarily at the footprints of newly heated loops where significant evaporation has not yet occurred. As a single loop evolves in time, no matter how strong the heating rate may become, the H(alpha) flux will diminish as the corona becomes denser and hence more effective at stopping the beam. This prediction leads to several observable consequences regarding the spatial and temporal signatures of the X-ray and H(alpha) radiation during flares.

Hawley, Suzanne L.↗

Electro-Active Polymer (EAP) Actuators for Planetary Applications

NASA is seeking to reduce the mass, size, consumed power, and cost of the instrumentation used in its future missions. An important element of many instruments and devices is the actuation mechanism and electroactive polymers (EAP) are offering an effective alternative to current actuators. In this study, two families of EAP materials were investigated, including bending ionomers and longitudinal electrostatically driven elastomers. These materials were demonstrated to effectively actuate manipulation devices and their performance is being enhanced in this on-going study. The recent observations are reported in this paper, include the operation of the bending-EAP at conditions that exceed the harsh environment on Mars, and identify the obstacles that its properties and characteristics are posing to using them as actuators. Analysis of the electrical characteristics of the ionomer EAP showed that it is a current driven material rather than voltage driven and the conductivity distribution on the surface of the material greatly influences the bending performance. An accurate equivalent circuit modeling of the ionomer EAP performance is essential for the design of effective drive electronics. The ionomer main limitations are the fact that it needs to be moist continuously and the process of electrolysis that takes place during activation. An effective coating technique using a sprayed polymer was developed extending its operation in air from a few minutes to about four months. The coating technique effectively forms the equivalent of a skin to protect the moisture content of the ionomer. In parallel to the development of the bending EAP, the development of computer control of actuated longitudinal EAP has been pursued. An EAP driven miniature robotic arm was constructed and it is controlled by a MATLAB code to drop and lift the arm and close and open EAP fingers of a 4-finger gripper. Keywords: Miniature Robotics, Electroactive Polymers, Electroactive Actuators, EAP Materials

Bar-Cohen, Y.↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Special purpose parallel computer architecture for real-time control and simulation in robotic applications

This is a real-time robotic controller and simulator which is a MIMD-SIMD parallel architecture for interfacing with an external host computer and providing a high degree of parallelism in computations for robotic control and simulation. It includes a host processor for receiving instructions from the external host computer and for transmitting answers to the external host computer. There are a plurality of SIMD microprocessors, each SIMD processor being a SIMD parallel processor capable of exploiting fine grain parallelism and further being able to operate asynchronously to form a MIMD architecture. Each SIMD processor comprises a SIMD architecture capable of performing two matrix-vector operations in parallel while fully exploiting parallelism in each operation. There is a system bus connecting the host processor to the plurality of SIMD microprocessors and a common clock providing a continuous sequence of clock pulses. There is also a ring structure interconnecting the plurality of SIMD microprocessors and connected to the clock for providing the clock pulses to the SIMD microprocessors and for providing a path for the flow of data and instructions between the SIMD microprocessors. The host processor includes logic for controlling the RRCS by interpreting instructions sent by the external host computer, decomposing the instructions into a series of computations to be performed by the SIMD microprocessors, using the system bus to distribute associated data among the SIMD microprocessors, and initiating activity of the SIMD microprocessors to perform the computations on the data by procedure call.

Fijany, Amir↗