Search NASA⌕ Search

SEARCH · Search NASA

Results for “Distributed Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61

Ozone transport during a cut-off low event studied in the frame of the TOASTE program

A study of ozone transfer to the troposphere has been performed during two phases of the evolution of a cut-off low using both ozone vertical profiles and objective analysis of the ECMWF to compute potential vorticity distributions and air mass trajectories. Ozone profiles were measured by a ground based lidar system at the Observatoire de Haute Provence (OHP, 43 deg 55 N, 5 deg 42 E). A stratospheric ozone transport into the troposphere has been observed during a tropopause fold which occurred at the beginning of the cut-off low formation and during the erosion phase of the cut-off low. From the estimate of the maximum ozone content transferred to the troposphere, both mechanisms have the same order of magnitude of influence on the ozone flux to the troposphere. On a time scale of a few days, the correlation is very good between the potential vorticity and the ozone time evolution in the vicinity of the upper level frontal system.

Ancellet, G.↗

Unstructured grids on SIMD torus machines

Unstructured grids lead to unstructured communication on distributed memory parallel computers, a problem that has been considered difficult. Here, we consider adaptive, offline communication routing for a SIMD processor grid. Our approach is empirical. We use large data sets drawn from supercomputing applications instead of an analytic model of communication load. The chief contribution of this paper is an experimental demonstration of the effectiveness of certain routing heuristics. Our routing algorithm is adaptive, nonminimal, and is generally designed to exploit locality. We have a parallel implementation of the router, and we report on its performance.

Bjorstad, Petter E.↗

Parallel Newton-Krylov-Schwarz algorithms for the transonic full potential equation

We study parallel two-level overlapping Schwarz algorithms for solving nonlinear finite element problems, in particular, for the full potential equation of aerodynamics discretized in two dimensions with bilinear elements. The overall algorithm, Newton-Krylov-Schwarz (NKS), employs an inexact finite-difference Newton method and a Krylov space iterative method, with a two-level overlapping Schwarz method as a preconditioner. We demonstrate that NKS, combined with a density upwinding continuation strategy for problems with weak shocks, is robust and, economical for this class of mixed elliptic-hyperbolic nonlinear partial differential equations, with proper specification of several parameters. We study upwinding parameters, inner convergence tolerance, coarse grid density, subdomain overlap, and the level of fill-in in the incomplete factorization, and report their effect on numerical convergence rate, overall execution time, and parallel efficiency on a distributed-memory parallel computer.

Cai, Xiao-Chuan↗

Visualization and Tracking of Parallel CFD Simulations

We describe a system for interactive visualization and tracking of a 3-D unsteady computational fluid dynamics (CFD) simulation on a parallel computer. CM/AVS, a distributed, parallel implementation of a visualization environment (AVS) runs on the CM-5 parallel supercomputer. A CFD solver is run as a CM/AVS module on the CM-5. Data communication between the solver, other parallel visualization modules, and a graphics workstation, which is running AVS, are handled by CM/AVS. Partitioning of the visualization task, between CM-5 and the workstation, can be done interactively in the visual programming environment provided by AVS. Flow solver parameters can also be altered by programmable interactive widgets. This system partially removes the requirement of storing large solution files at frequent time steps, a characteristic of the traditional 'simulate (yields) store (yields) visualize' post-processing approach.

Vaziri, Arsi↗

Load Balancing Unstructured Adaptive Grids for CFD Problems

Mesh adaption is a powerful tool for efficient unstructured-grid computations but causes load imbalance among processors on a parallel machine. A dynamic load balancing method is presented that balances the workload across all processors with a global view. After each parallel tetrahedral mesh adaption, the method first determines if the new mesh is sufficiently unbalanced to warrant a repartitioning. If so, the adapted mesh is repartitioned, with new partitions assigned to processors so that the redistribution cost is minimized. The new partitions are accepted only if the remapping cost is compensated by the improved load balance. Results indicate that this strategy is effective for large-scale scientific computations on distributed-memory multiprocessors.

Biswas, Rupak↗

Balancing Contention and Synchronization on the Intel Paragon

The Intel Paragon is a mesh-connected distributed memory parallel computer. It uses an oblivious and deterministic message routing algorithm: this permits us to develop highly optimized schedules for frequently needed communication patterns. The complete exchange is one such pattern. Several approaches are available for carrying it out on the mesh. We study an algorithm developed by Scott. This algorithm assumes that a communication link can carry one message at a time and that a node can only transmit one message at a time. It requires global synchronization to enforce a schedule of transmissions. Unfortunately global synchronization has substantial overhead on the Paragon. At the same time the powerful interconnection mechanism of this machine permits 2 or 3 messages to share a communication link with minor overhead. It can also overlap multiple message transmission from the same node to some extent. We develop a generalization of Scott's algorithm that executes complete exchange with a prescribed contention. Schedules that incur greater contention require fewer synchronization steps. This permits us to tradeoff contention against synchronization overhead. We describe the performance of this algorithm and compare it with Scott's original algorithm as well as with a naive algorithm that does not take interconnection structure into account. The Bounded contention algorithm is always better than Scott's algorithm and outperforms the naive algorithm for all but the smallest message sizes. The naive algorithm fails to work on meshes larger than 12 x 12. These results show that due consideration of processor interconnect and machine performance parameters is necessary to obtain peak performance from the Paragon and its successor mesh machines.

Bokhari, Shahid H.↗

A Parallel Pipelined Renderer for the Time-Varying Volume Data

This paper presents a strategy for efficiently rendering time-varying volume data sets on a distributed-memory parallel computer. Time-varying volume data take large storage space and visualizing them requires reading large files continuously or periodically throughout the course of the visualization process. Instead of using all the processors to collectively render one volume at a time, a pipelined rendering process is formed by partitioning processors into groups to render multiple volumes concurrently. In this way, the overall rendering time may be greatly reduced because the pipelined rendering tasks are overlapped with the I/O required to load each volume into a group of processors; moreover, parallelization overhead may be reduced as a result of partitioning the processors. We modify an existing parallel volume renderer to exploit various levels of rendering parallelism and to study how the partitioning of processors may lead to optimal rendering performance. Two factors which are important to the overall execution time are re-source utilization efficiency and pipeline startup latency. The optimal partitioning configuration is the one that balances these two factors. Tests on Intel Paragon computers show that in general optimal partitionings do exist for a given rendering task and result in 40-50% saving in overall rendering time.

Chiueh, Tzi-Cker↗

Some Examples of the Applications of the Transonic and Supersonic Area Rules to the Prediction of Wave Drag

The experimental wave drags of bodies and wing-body combinations over a wide range of Mach numbers are compared with the computed drags utilizing a 24-term Fourier series application of the supersonic area rule and with the results of equivalent-body tests. The results indicate that the equivalent-body technique provides a good method for predicting the wave drag of certain wing-body combinations at and below a Mach number of 1. At Mach numbers greater than 1, the equivalent-body wave drags can be misleading. The wave drags computed using the supersonic area rule are shown to be in best agreement with the experimental results for configurations employing the thinnest wings. The wave drags for the bodies of revolution presented in this report are predicted to a greater degree of accuracy by using the frontal projections of oblique areas than by using normal areas. A rapid method of computing wing area distributions and area-distribution slopes is given in an appendix.

Nelson, Robert L.↗

Aerodynamic Shape Optimization Using A Combined Distributed/Shared Memory Paradigm

Current parallel computational approaches involve distributed and shared memory paradigms. In the distributed memory paradigm, each processor has its own independent memory. Message passing typically uses a function library such as MPI or PVM. In the shared memory paradigm, such as that used on the SGI Origin 2000 machine, compiler directives are used to instruct the compiler to schedule multiple threads to perform calculations. In this paradigm, it must be assured that processors (threads) do not simultaneously access regions of memory in such away that errors would occur. This paper utilizes the latest version of the SGI MPI function library to combine the two parallelization paradigms to perform aerodynamic shape optimization of a generic wing/body.

Cheung, Samson↗

A Java-Enabled Interactive Graphical Gas Turbine Propulsion System Simulator

This paper describes a gas turbine simulation system which utilizes the newly developed Java language environment software system. The system provides an interactive graphical environment which allows the quick and efficient construction and analysis of arbitrary gas turbine propulsion systems. The simulation system couples a graphical user interface, developed using the Java Abstract Window Toolkit, and a transient, space- averaged, aero-thermodynamic gas turbine analysis method, both entirely coded in the Java language. The combined package provides analytical, graphical and data management tools which allow the user to construct and control engine simulations by manipulating graphical objects on the computer display screen. Distributed simulations, including parallel processing and distributed database access across the Internet and World-Wide Web (WWW), are made possible through services provided by the Java environment.

Reed, John A.↗

Assimilation of Surface Temperature in Land Surface Models

Hydrological models have been calibrated and validated using catchment streamflows. However, using a point measurement does not guarantee correct spatial distribution of model computed heat fluxes, soil moisture and surface temperatures. With the advent of satellites in the late 70s, surface temperature is being measured two to four times a day from various satellite sensors and different platforms. The purpose of this paper is to demonstrate use of satellite surface temperature in (a) validation of model computed surface temperatures and (b) assimilation of satellite surface temperatures into a hydrological model in order to improve the prediction accuracy of soil moistures and heat fluxes. The assimilation is carried out by comparing the satellite and the model produced surface temperatures and setting the "true"temperature midway between the two values. Based on this "true" surface temperature, the physical relationships of water and energy balance are used to reset the other variables. This is a case of nudging the water and energy balance variables so that they are consistent with each other and the true" surface temperature. The potential of this assimilation scheme is demonstrated in the form of various experiments that highlight the various aspects. This study is carried over the Red-Arkansas basin in the southern United States (a 5 deg X 10 deg area) over a time period of a year (August 1987 - July 1988). The land surface hydrological model is run on an hourly time step. The results show that satellite surface temperature assimilation improves the accuracy of the computed surface soil moisture remarkably.

Lakshmi, Venkataraman↗

Evaluation of SAGE II and Balloon-Borne Stratospheric Aerosol Measurements: Evaluation of Aerosol Measurements from SAGE II, HALOE, and Balloonborne Optical Particle Counters

Stratospheric aerosol measurements from the University of Wyoming balloonborne optical particle counters (OPCs), the Stratospheric Aerosol and Gas Experiment (SAGE) II, and the Halogen Occultation Experiment (HALOE) were compared in the period 1982-2000, when measurements were available. The OPCs measure aerosol size distributions, and HALOE multiwavelength (2.45-5.26 micrometers) extinction measurements can be used to retrieve aerosol size distributions. Aerosol extinctions at the SAGE II wavelengths (0.386-1.02 micrometers) were computed from these size distributions and compared to SAGE II measurements. In addition, surface areas derived from all three experiments were compared. While the overall impression from these results is encouraging, the agreement can change with latitude, altitude, time, and parameter. In the broadest sense, these comparisons fall into two categories: high aerosol loading (volcanic periods) and low aerosol loading (background periods and altitudes above 25 km). When the aerosol amount was low, SAGE II and HALOE extinctions were higher than the OPC estimates, while the SAGE II surface areas were lower than HALOE and the OPCS. Under high loading conditions all three instruments mutually agree to within 50%.

Hervig, Mark↗

Simple LDAP Schemas for Grid Monitoring

The purpose of this document is to provide an initial definition of the data we need in a directory service or database so that we can implement a performance monitoring testbed. To begin with, this document describes how to represent producers of events and event schemes. The representation of producers is simple and does not contain information such as who has access to the events and what protocols can be used to access the events. In the future, we will define how to describe consumers of events and add details to our representations. A popular choice for a directory service or database for grid computing is a distributed directory service that is accessed using the Lightweight Directory Access Protocol (LDAP). This document uses LDAP terminology, schemes, and formats to represent the directory service schemes.

Smith, Warren↗

New NAS Parallel Benchmarks Results

NPB2 (NAS (NASA Advanced Supercomputing) Parallel Benchmarks 2) is an implementation, based on Fortran and the MPI (message passing interface) message passing standard, of the original NAS Parallel Benchmark specifications. NPB2 programs are run with little or no tuning, in contrast to NPB vendor implementations, which are highly optimized for specific architectures. NPB2 results complement, rather than replace, NPB results. Because they have not been optimized by vendors, NPB2 implementations approximate the performance a typical user can expect for a portable parallel program on distributed memory parallel computers. Together these results provide an insightful comparison of the real-world performance of high-performance computers. New NPB2 features: New implementation (CG), new workstation class problem sizes, new serial sample versions, more performance statistics.

Yarrow, Maurice↗

Accounting and Accountability for Distributed and Grid Systems

While the advent of distributed and grid computing systems will open new opportunities for scientific exploration, the reality of such implementations could prove to be a system administrator's nightmare. A lot of effort is being spent on identifying and resolving the obvious problems of security, scheduling, authentication and authorization. Lurking in the background, though, are the largely unaddressed issues of accountability and usage accounting: (1) mapping resource usage to resource users; (2) defining usage economies or methods for resource exchange; (3) describing implementation standards that minimize and compartmentalize the tasks required for a site to participate in a grid.

Thigpen, William↗

The Frequency Detuning Correction and the Asymmetry of Line Shapes: The Far Wings of H2O-H2O

A far-wing line shape theory which satisfies the detailed balance principle is applied to the H2O-H2O system. Within this formalism, two line shapes are introduced, corresponding to band-averages over the positive and negative resonance lines, respectively. Using the coordinate representation, the two line shapes can be obtained by evaluating 11-dimensional integrations whose integrands are a product of two factors. One depends on the interaction between the two molecules and is easy to evaluate. The other contains the density matrix of the system and is expressed as a product of two 3-dimensional distributions associated with the density matrices of the absorber and the perturber molecule, respectively. If most of the populated states are included in the averaging process, to obtain these distributions requires extensive computer CPU time, but only have to be computed once for a given temperature. The 11-dimensional integrations are evaluated using the Monte Carlo method, and in order to reduce the variance, the integration variables are chosen such that the sensitivity of the integrands on them is clearly distinguished.

Ma, Q.↗

Lightning Return-Stroke Current Waveforms Aloft, From Measured Field Change, Current, and Channel Geometry

Direct current measurements are available near the attachment point from both natural cloud-to-ground lightning and rocket-triggered lightning, but little is known about the rise time and peak amplitude of return-stroke currents aloft. We present, as functions of height, current amplitudes, rise times, and effective propagation velocities that have been estimated with a novel remote-sensing technique from data on 24 subsequent return strokes in six different lightning flashes that were triggering at the NASA Kennedy Space Center, FL, during 1987. The unique feature of this data set is the stereo pairs of still photographs, from which three-dimensional channel geometries were determined previously. This has permitted us to calculate the fine structure of the electric-field-change (E) waveforms produced by these strokes, using the current waveforms measured at the channel base together with physically reasonable assumptions about the current distributions aloft. The computed waveforms have been compared with observed E waveforms from the same strokes, and our assumptions have been adjusted to maximize agreement. In spite of the non-uniqueness of solutions derived by this technique, several conclusions seem inescapable: 1) The effective propagation speed of the current up the channel is usually significantly (but not unreasonably) faster than the two-dimensional velocity measured by a streak camera for 14 of these strokes. 2) Given the deduced propagation speed, the peak amplitude of the current waveform often must decrease dramatically with height to prevent the electric field from being over-predicted. 3) The rise time of the current wave front must always increase rapidly with height in order to keep the fine structure of the calculated field consistent with the observations.

Willett, J. C.↗

The Compressible Potential Flow Past Elliptic Symmetrical Cylinders at Zero Angle of Attack and with No Circulation

For the tunnel corrections of compressible flows those profiles are of interest for which at least the second approximation of the Janzen-Rayleigh method can be applied in closed form. One such case is presented by certain elliptical symmetrical cylinders located in the center of a tunnel with fixed walls and whose maximum velocity, incompressible, is twice the velocity of flow. In the numerical solution the maximum velocity at the profile and the tunnel wall as well as the entry of sonic velocity is computed. The velocity distribution past the contour and in the minimum cross section at various Mach numbers is illustrated on a worked out-example.

Hantzsche, W.↗