Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Directions in parallel programming: HPF, shared virtual memory and object parallelism in pC++

Fortran and C++ are the dominant programming languages used in scientific computation. Consequently, extensions to these languages are the most popular for programming massively parallel computers. We discuss two such approaches to parallel Fortran and one approach to C++. The High Performance Fortran Forum has designed HPF with the intent of supporting data parallelism on Fortran 90 applications. HPF works by asking the user to help the compiler distribute and align the data structures with the distributed memory modules in the system. Fortran-S takes a different approach in which the data distribution is managed by the operating system and the user provides annotations to indicate parallel control regions. In the case of C++, we look at pC++ which is based on a concurrent aggregate parallel model.

Bodin, Francois↗

Use of networked workstations for parallel nonlinear structural dynamic simulations of rotating bladed-disk assemblies

The principal objective of this research is to investigate, develop and demonstrate coarse-grained, parallel-processing strategies for nonlinear dynamic simulations for rotating bladed-disk assemblies. The parallel -processing strategies addressed include numerical algorithms for parallel nonlinear solutions and techniques to effect load balancing among processors. The parallel environment employed is a distributed-memory, coarse-grained one consisting of networked workstations. A parallel explicit time integration method has been implemented for transient nonlinear solutions of rotationg bladed-disk assemblies. Automatic domain partitioning techniques have been investigated for load balancing among processors. Advanced computing environments, data structures and interactive computer graphics all contribute to an integrated parallel finite element analysis system to facilitate more efficient and powerful dynamic simulations.

Hsieh, Shang-Hsien↗

Impact of Different Thermal Gradients on the Dynamics of Cylindrical Lithium-ion Cells Subject to Accelerated Aging and on Module Performance

This study investigates the impacts of applying different thermal gradient patterns to cylindrical lithium-ion cells in a module on cell dynamics (temperatures, current flows, state of charge), module performance (evolution of resistance, capacity, and energy versus cycle number), and module lifetime. The thermal gradients were generated using cooling plates (CPs) with three different flow-field designs, namely, straight, perpendicular, and U-turn. The study uses computational fluid dynamics (CFD), the pseudo-two-dimensional (P2D) battery model, capacity loss and increased impedance due to the growth of a solid-electrolyte-interphase, and the electric current distribution from module terminals to cells that depends on the series-parallel electrical connections among the cells. The impact of the thermal gradient (resulting from the CP designs) on the variability in resistance, current, state of charge, and voltage among the cells was analyzed and linked to differences in the module's performance. Applying a thermal gradient to parallel-connected strings of series-connected cells led to variation in the current through each parallel string and an imbalance in the voltage of series-connected cells. Module performance is poorer when the thermal gradient causes a voltage imbalance than when it causes a current imbalance. Module performance becomes the worst when both current variation and voltage imbalance happen together. For instance, the module's lifetime (estimated as reaching 80% of its initial capacity) varied by 5% to 17.5%, depending on the magnitude and pattern of the imposed thermal gradient. As the relative orientation between thermal gradients and cells' electrical connectivity influences the module's performance, appropriate consideration should be given to the choice of the CP, especially if large thermal gradients are allowed.

Battery thermal management↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗

Velocity-space synthesis of ISEE-1 measurements of the three dimensional electron distribution function

A computer package which produces contour plots of the three dimensional electron distribution function measured by an electron spectrometer aboard ISEE-1 is described. Examples of the contour plots and an explanation of how to use the program, including the necessary computer code for running the program on the GSFC 360/91 computer is presented. The method by which the discrete measurements of the distribution function, given by points on the four dimensional surface are synthesized into a smooth surface in a three dimensional space which can be contoured is described. The velocity components are parallel and perpendicular to the magnetic field, respectively, in the proper frame of the electrons.

Fitzenreiter, R. J.↗

Coupling of newborn ions to the solar wind by electromagnetic instabilities and their interaction with the bow shock

The process by which the solar wind assimilates newly ionized atoms is important for understanding the presence of planetary or interstellar helium in the solar wind, the dynamics of the Active Magnetospheric Particle Tracer Explorers (AMPTE) lithium releases in front of the earth's bow shock, and the formation of cometary tails. In this paper is examined how newborn ions can be coupled to the solar wind in the direction parallel to the magnetic field by means of electromagnetic instabilities driven by the distribution of newborn ions. The linear properties of three instabilities are analyzed and compared with numerical solutions of the linear dispersion equation, while their nonlinear behavior is followed by means of computer simulation to obtain the characteristic time for the pickup process. With a primary emphasis on the AMPTE lithiuim releases, various degrees of realism are introduced into the calculations to model the upstream conditions and the intersection of the lithium with the bow shock. It is shown that a time-dependent shock model is needed to correctly reproduce the amount of lithium which is transmitted through the shock and that the resulting lithium ion distribution is still likely to be subject to the same type of instabilities in the magnetosheath. Applications of these results to comets, in particular the artificial comet expected to be generated by the AMPTE barium release in the magnetosheath, is also briefly discussed.

Winske, D.↗

A solar chromosphere and spicule model based on far-infrared limb observations

Techniques developed for LTE radiative transfer problems in a rough atmosphere were used to compute a model chromosphere containing spicules consistent with high-resolution solar limb observations from 100 microns to 2.6 mm. The model consists of a smooth, plane-parallel temperature minimum region extending from the photosphere to a height of 1000 km and randomly distributed cylindrical spicules above this height. It is found that the observed limb brightness profiles are well fitted by spicules with electron temperatures on the order of 7000 K.

Braun, D.↗

Parallel Computation of Unsteady Flows on a Network of Workstations

Parallel computation of unsteady flows requires significant computational resources. The utilization of a network of workstations seems an efficient solution to the problem where large problems can be treated at a reasonable cost. This approach requires the solution of several problems: 1) the partitioning and distribution of the problem over a network of workstation, 2) efficient communication tools, 3) managing the system efficiently for a given problem. Of course, there is the question of the efficiency of any given numerical algorithm to such a computing system. NPARC code was chosen as a sample for the application. For the explicit version of the NPARC code both two- and three-dimensional problems were studied. Again both steady and unsteady problems were investigated. The issues studied as a part of the research program were: 1) how to distribute the data between the workstations, 2) how to compute and how to communicate at each node efficiently, 3) how to balance the load distribution. In the following, a summary of these activities is presented. Details of the work have been presented and published as referenced.

Source record↗

The impact of non-local parallel electron transport on plasma-impurity reaction rates in tokamak scrape-off layer plasmas

Abstract Plasma-impurity reaction rates are a crucial part of modelling tokamak scrape-off layer (SOL) plasmas. To avoid calculating the full set of rates for the large number of important processes involved, a set of effective rates are typically derived which assume Maxwellian electrons. However, non-local parallel electron transport may result in non-Maxwellian electrons, particularly close to divertor targets. Here, the validity of using Maxwellian-averaged rates in this context is investigated by computing the full set of rate equations for a fixed plasma background from kinetic and fluid SOL simulations. We consider the effect of the electron distribution as well as the impact of the electron transport model on plasma profiles. Results are presented for lithium, beryllium, carbon, nitrogen, neon and argon. It is found that electron distributions with enhanced high-energy tails can result in significant modifications to the ionisation balance and radiative power loss rates from excitation, on the order of 50%–75% for the latter. Fluid electron models with Spitzer-Härm or flux-limited Spitzer-Härm thermal conductivity, combined with Maxwellian electrons for rate calculations, can increase or decrease this error, depending on the impurity species and plasma conditions. Based on these results, we also discuss some approaches to experimentally observing non-local electron transport in SOL plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

State-of-the-art Space Telescope Digicon performance data

The Digicon has been chosen as the detector for the High Resolution Spectrograph and the Faint Object Spectrograph of the Space Telescope. Both tubes are 512 channel, parallel-output devices and feature CsTe photocathodes on MgF2 faceplates. Using a computer-assisted test facility, the tubes have been characterized with respect to diode array performance, photocathode response (1100-9000 A), and imaging capability. Data are presented on diode dark current and capacitance distributions, pulse height resolution, photocathode quantum efficiency, uniformity and blemishes, dark count rate, distortion, resolution, and crosstalk.

Ginaven, R. O.↗

Asynchronous interactive control systems

A class of interactive control systems is derived by generalizing interactive manipulator control systems. The general structural properties of such systems are discussed and an appropriate general software implementation is proposed. This is based on the fact that tasks of interactive control systems can be represented as a network of a finite set of actions which have specific operational characteristics and specific resource requirements, and which are of limited duration. This has enabled the decomposition of the overall control algorithm into a set of subalgorithms, called subcontrollers, which can operate simultaneously and asynchronously. Coordinate transformations of sensor feedback data and actuator set-points have enabled the further simplification of the subcontrollers and have reduced their conflicting resource requirements. The modules of the decomposed control system are implemented as parallel processes with disjoint memory space communicating only by I/O. The synchronization mechanisms for dynamic resource allocation among subcontrollers and other synchronization mechanisms are also discussed in this paper. Such a software organization is suitable for the general form of multiprocessing using computer networks with distributed storage.

Vuskovic, M. I.↗

Ponderomotive effects on distributions of O(+) ions in the auroral zone

Test particle calculations are used to compute the effects of gravity and ponderomotive acceleration by shear Alfven wave oscillations on the distribution function of O(+) ions along auroral field lines, assuming an ionospheric Maxwellian source of the ions at 2000 km altitude with approximately 0.5 eV of thermal energy in the parallel component of velocity. The electric field model corresponds to a standing wave oscillation with a frequency approximately 1 Hz in the azimuthal direction superimposed on the background dipole field, in which the wave amplitude is either increasing or decreasing in time. The electric field is taken to be primarily in the perpendicular direction. The time varying wave produces broad distributions with widths of 2 to 10 times the initial 0.5-eV thermal energy of the Maxwellian source, and the density and flux of upward going O(+) ions at one Earth radius are both enhanced in this model. The oxygen ion distribution functions at 1 R(sub E) altitude resulting from interaction with waves whose amplitudes are increasing in time have a more gradual lower energy cutoff than do the distribution functions resulting from decaying waves. The high-energy part of the distribution functions in growing waves reflects the temperature of the Maxwellian source, while the high-energy part of the distributions resulting from decaying waves steepens with time, independent of the source temperature.

Witt, E.↗

GSRP/David Marshall: Fully Automated Cartesian Grid CFD Application for MDO in High Speed Flows

With the renewed interest in Cartesian gridding methodologies for the ease and speed of gridding complex geometries in addition to the simplicity of the control volumes used in the computations, it has become important to investigate ways of extending the existing Cartesian grid solver functionalities. This includes developing methods of modeling the viscous effects in order to utilize Cartesian grids solvers for accurate drag predictions and addressing the issues related to the distributed memory parallelization of Cartesian solvers. This research presents advances in two areas of interest in Cartesian grid solvers, viscous effects modeling and MPI parallelization. The development of viscous effects modeling using solely Cartesian grids has been hampered by the widely varying control volume sizes associated with the mesh refinement and the cut cells associated with the solid surface. This problem is being addressed by using physically based modeling techniques to update the state vectors of the cut cells and removing them from the finite volume integration scheme. This work is performed on a new Cartesian grid solver, NASCART-GT, with modifications to its cut cell functionality. The development of MPI parallelization addresses issues associated with utilizing Cartesian solvers on distributed memory parallel environments. This work is performed on an existing Cartesian grid solver, CART3D, with modifications to its parallelization methodology.

Source record↗

Supercomputing Aspects for Simulating Incompressible Flow

The primary objective of this research is to support the design of liquid rocket systems for the Advanced Space Transportation System. Since the space launch systems in the near future are likely to rely on liquid rocket engines, increasing the efficiency and reliability of the engine components is an important task. One of the major problems in the liquid rocket engine is to understand fluid dynamics of fuel and oxidizer flows from the fuel tank to plume. Understanding the flow through the entire turbo-pump geometry through numerical simulation will be of significant value toward design. One of the milestones of this effort is to develop, apply and demonstrate the capability and accuracy of 3D CFD methods as efficient design analysis tools on high performance computer platforms. The development of the Message Passage Interface (MPI) and Multi Level Parallel (MLP) versions of the INS3D code is currently underway. The serial version of INS3D code is a multidimensional incompressible Navier-Stokes solver based on overset grid technology, INS3D-MPI is based on the explicit massage-passing interface across processors and is primarily suited for distributed memory systems. INS3D-MLP is based on multi-level parallel method and is suitable for distributed-shared memory systems. For the entire turbo-pump simulations, moving boundary capability and efficient time-accurate integration methods are built in the flow solver, To handle the geometric complexity and moving boundary problems, an overset grid scheme is incorporated with the solver so that new connectivity data will be obtained at each time step. The Chimera overlapped grid scheme allows subdomains move relative to each other, and provides a great flexibility when the boundary movement creates large displacements. Two numerical procedures, one based on artificial compressibility method and the other pressure projection method, are outlined for obtaining time-accurate solutions of the incompressible Navier-Stokes equations. The performance of the two methods is compared by obtaining unsteady solutions for the evolution of twin vortices behind a flat plate. Calculated results are compared with experimental and other numerical results. For an unsteady flow, which requires small physical time step, the pressure projection method was found to be computationally efficient since it does not require any subiteration procedure. It was observed that the artificial compressibility method requires a fast convergence scheme at each physical time step in order to satisfy the incompressibility condition. This was obtained by using a GMRES-ILU(0) solver in present computations. When a line-relaxation scheme was used, the time accuracy was degraded and time-accurate computations became very expensive.

Kwak, Dochan↗

Globalized Newton-Krylov-Schwarz Algorithms and Software for Parallel Implicit CFD

Implicit solution methods are important in applications modeled by PDEs with disparate temporal and spatial scales. Because such applications require high resolution with reasonable turnaround, "routine" parallelization is essential. The pseudo-transient matrix-free Newton-Krylov-Schwarz (Psi-NKS) algorithmic framework is presented as an answer. We show that, for the classical problem of three-dimensional transonic Euler flow about an M6 wing, Psi-NKS can simultaneously deliver: globalized, asymptotically rapid convergence through adaptive pseudo- transient continuation and Newton's method-, reasonable parallelizability for an implicit method through deferred synchronization and favorable communication-to-computation scaling in the Krylov linear solver; and high per- processor performance through attention to distributed memory and cache locality, especially through the Schwarz preconditioner. Two discouraging features of Psi-NKS methods are their sensitivity to the coding of the underlying PDE discretization and the large number of parameters that must be selected to govern convergence. We therefore distill several recommendations from our experience and from our reading of the literature on various algorithmic components of Psi-NKS, and we describe a freely available, MPI-based portable parallel software implementation of the solver employed here.

Gropp, W. D.↗

Consequences of using nonlinear particle trajectories to compute spatial diffusion coefficients

In a study of cosmic ray propagation in interstellar and interplanetary space, a perturbed orbit resonant scattering theory for pitch angle diffusion in a slab model of magnetostatic turbulence is slightly generalized and used to compute the diffusion coefficient for spatial propagation parallel to the mean magnetic field. This diffusion coefficient has been useful for describing the solar modulation of the galactic cosmic rays, and for explaining the diffusive phase in solar flares in which the initial anisotropy of the particle distribution decays to isotropy.

Goldstein, M. L.↗

Latency Hiding in Dynamic Partitioning and Load Balancing of Grid Computing Applications

The Information Power Grid (IPG) concept developed by NASA is aimed to provide a metacomputing platform for large-scale distributed computations, by hiding the intricacies of highly heterogeneous environment and yet maintaining adequate security. In this paper, we propose a latency-tolerant partitioning scheme that dynamically balances processor workloads on the.IPG, and minimizes data movement and runtime communication. By simulating an unsteady adaptive mesh application on a wide area network, we study the performance of our load balancer under the Globus environment. The number of IPG nodes, the number of processors per node, and the interconnected speeds are parameterized to derive conditions under which the IPG would be suitable for parallel distributed processing of such applications. Experimental results demonstrate that effective solution are achieved when the IPG nodes are connected by a high-speed asynchronous interconnection network.

Das, Sajal K.↗

Applications of Parallel-Element, Embedded Mesh-Cap Acoustic Liner Concepts

This study explores progress achieved with 2DOF, 3DOF, and MDOF acoustic liners constructed with mesh caps embedded within a honeycomb core. These liner configurations offer potential for broadband noise reduction, and are suitable for conventional aircraft implementation. Samples for each configuration are tested in the NASA normal incidence tube and grazing flow impedance tube, with and without a wire mesh facesheet. Impedances based on these measured data compare favorably with those predicted using a transmission line impedance prediction model. Predicted impedances are then used as input for an aeroacoustic propagation code to compute axial acoustic pressure distributions in the grazing flow tube. These predicted distributions compare favorably with the corresponding measured distributions at frequencies away from the frequency of peak attenuation, but suffer slight degradation for frequencies very near the peak attenuation frequency, where the predicted results are sensitive to input impedance changes. As expected, the noise reduction frequency range increases as more degrees of freedom are included. Although the specific results achieved herein may differ from those that would be achieved with other 2DOF, 3DOF, and MDOF liners, this comparison highlights some of the key features that can be exploited in the design of parallel-element, embedded mesh-cap liners.

Jones, M. G.↗