Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60

A real-time diagnostic and performance monitor for UNIX

There are now over one million UNIX sites and the pace at which new installations are added is steadily increasing. Along with this increase, comes a need to develop simple efficient, effective and adaptable ways of simultaneously collecting real-time diagnostic and performance data. This need exists because distributed systems can give rise to complex failure situations that are often un-identifiable with single-machine diagnostic software. The simultaneous collection of error and performance data is also important for research in failure prediction and error/performance studies. This paper introduces a portable method to concurrently collect real-time diagnostic and performance data on a distributed UNIX system. The combined diagnostic/performance data collection is implemented on a distributed multi-computer system using SUN4's as servers. The approach uses existing UNIX system facilities to gather system dependability information such as error and crash reports. In addition, performance data such as CPU utilization, disk usage, I/O transfer rate and network contention is also collected. In the future, the collected data will be used to identify dependability bottlenecks and to analyze the impact of failures on system performance.

Dong, Hongchao↗

National Space Science Data Center data archive and distribution service (NDADS) automated retrieval mail system user's guide

The National Space Science Data Center (NSSDC) has developed an automated data retrieval request service utilizing our Data Archive and Distribution Service (NDADS) computer system. NDADS currently has selected project data written to optical disk platters with the disks residing in a robotic 'jukebox' near-line environment. This allows for rapid and automated access to the data with no staff intervention required. There are also automated help information and user services available that can be accessed. The request system permits an average-size data request to be completed within minutes of the request being sent to NSSDC. A mail message, in the format described in this document, retrieves the data and can send it to a remote site. Also listed in this document are the data currently available.

Perry, Charleen M.↗

The velocity dispersion of the caustic network due to random motion of individual stars in the lensing galaxy

We present a method of computing the velocity distribution of the caustic network due to the random motion of stars in the lensing galaxy. This method is illustrated on the example of the two-point mass lens and then applied to a large sample of stars. We conclude that the proper motion of the stars increases significantly the frequency of the high magnification events in comparison with a static lens configuration with constant stream or constant bulk velocity. The stream velocity is the velocity of the star field relative to the global bulk velocity of the galaxy. We show that the global bulk and the stream velocity of the star field have to be considered separately for any microlensing situation. The higher the surface mass density of the stars in the lensing galaxy, the higher the influence of proper motion of stars on the statistics of high magnification events. The influence of a Gaussian velocity distribution of the stars in the lensing galaxy compared with a constant stream velocity of the stars increases the number of high magnification events by a factor 1.30 +/- 0.06 for a normalized surface density of the stars cr = 0.1 and by a factor 1.7 +/- 0.1 for sigma = 0.5. This means that for some microlensing situations the proper motion of the stars in a lensing galaxy has to be considered for exact microlensing predictions.

Kundik, Tomislav↗

Generating local addresses and communication sets for data-parallel programs

Generating local addresses and communication sets is an important issue in distributed-memory implementations of data-parallel languages such as High Performance FORTRAN. We show that, for an array A affinely aligned to a template that is distributed across p processors with a cyclic(k) distribution and a computation involving the regular section A(l:h:s), the local memory access sequence for any processor is characterized by a finite state machine of at most k states. We present fast algorithms for computing the essential information about these state machines, and extend the framework to handle multidimensional arrays. We also show how to generate communication sets using the state machine approach. Performance results show that this solution requires very little run-time overhead and acceptable preprocessing time.

Chatterjee, Siddhartha↗

Communications among elements of a space construction ensemble

Space construction projects will require careful coordination between managers, designers, manufacturers, operators, astronauts, and robots with large volumes of information of varying resolution, timeliness, and accuracy flowing between the distributed participants over computer communications networks. Within the CSC Operations Branch, we are researching the requirements and options for such communications. Based on our work to date, we feel that communications standards being developed by the International Standards Organization, the CCITT, and other groups can be applied to space construction. We are currently studying in depth how such standards can be used to communicate with robots and automated construction equipment used in a space project. Specifically, we are looking at how the Manufacturing Automation Protocol (MAP) and the Manufacturing Message Specification (MMS), which tie together computers and machines in automated factories, might be applied to space construction projects. Together with our CSC industrial partner Computer Technology Associates, we are developing a MAP/MMS companion standard for space construction and we will produce software to allow the MAP/MMS protocol to be used in our CSC operations testbed.

Davis, Randal L.↗

Generating local addresses and communication sets for data-parallel programs

Generating local addresses and communication sets is an important issue in distributed-memory implementations of data-parallel languages such as High Performance Fortran. We show that for an array A affinely aligned to a template that is distributed across p processors with a cyclic(k) distribution, and a computation involving the regular section A, the local memory access sequence for any processor is characterized by a finite state machine of at most k states. We present fast algorithms for computing the essential information about these state machines, and extend the framework to handle multidimensional arrays. We also show how to generate communication sets using the state machine approach. Performance results show that this solution requires very little runtime overhead and acceptable preprocessing time.

Chatterjee, Siddhartha↗

Ozone transport during a cut-off low event studied in the frame of the TOASTE program

A study of ozone transfer to the troposphere has been performed during two phases of the evolution of a cut-off low using both ozone vertical profiles and objective analysis of the ECMWF to compute potential vorticity distributions and air mass trajectories. Ozone profiles were measured by a ground based lidar system at the Observatoire de Haute Provence (OHP, 43 deg 55 N, 5 deg 42 E). A stratospheric ozone transport into the troposphere has been observed during a tropopause fold which occurred at the beginning of the cut-off low formation and during the erosion phase of the cut-off low. From the estimate of the maximum ozone content transferred to the troposphere, both mechanisms have the same order of magnitude of influence on the ozone flux to the troposphere. On a time scale of a few days, the correlation is very good between the potential vorticity and the ozone time evolution in the vicinity of the upper level frontal system.

Ancellet, G.↗

Unstructured grids on SIMD torus machines

Unstructured grids lead to unstructured communication on distributed memory parallel computers, a problem that has been considered difficult. Here, we consider adaptive, offline communication routing for a SIMD processor grid. Our approach is empirical. We use large data sets drawn from supercomputing applications instead of an analytic model of communication load. The chief contribution of this paper is an experimental demonstration of the effectiveness of certain routing heuristics. Our routing algorithm is adaptive, nonminimal, and is generally designed to exploit locality. We have a parallel implementation of the router, and we report on its performance.

Bjorstad, Petter E.↗

Parallel Newton-Krylov-Schwarz algorithms for the transonic full potential equation

We study parallel two-level overlapping Schwarz algorithms for solving nonlinear finite element problems, in particular, for the full potential equation of aerodynamics discretized in two dimensions with bilinear elements. The overall algorithm, Newton-Krylov-Schwarz (NKS), employs an inexact finite-difference Newton method and a Krylov space iterative method, with a two-level overlapping Schwarz method as a preconditioner. We demonstrate that NKS, combined with a density upwinding continuation strategy for problems with weak shocks, is robust and, economical for this class of mixed elliptic-hyperbolic nonlinear partial differential equations, with proper specification of several parameters. We study upwinding parameters, inner convergence tolerance, coarse grid density, subdomain overlap, and the level of fill-in in the incomplete factorization, and report their effect on numerical convergence rate, overall execution time, and parallel efficiency on a distributed-memory parallel computer.

Cai, Xiao-Chuan↗

Visualization and Tracking of Parallel CFD Simulations

We describe a system for interactive visualization and tracking of a 3-D unsteady computational fluid dynamics (CFD) simulation on a parallel computer. CM/AVS, a distributed, parallel implementation of a visualization environment (AVS) runs on the CM-5 parallel supercomputer. A CFD solver is run as a CM/AVS module on the CM-5. Data communication between the solver, other parallel visualization modules, and a graphics workstation, which is running AVS, are handled by CM/AVS. Partitioning of the visualization task, between CM-5 and the workstation, can be done interactively in the visual programming environment provided by AVS. Flow solver parameters can also be altered by programmable interactive widgets. This system partially removes the requirement of storing large solution files at frequent time steps, a characteristic of the traditional 'simulate (yields) store (yields) visualize' post-processing approach.

Vaziri, Arsi↗

Load Balancing Unstructured Adaptive Grids for CFD Problems

Mesh adaption is a powerful tool for efficient unstructured-grid computations but causes load imbalance among processors on a parallel machine. A dynamic load balancing method is presented that balances the workload across all processors with a global view. After each parallel tetrahedral mesh adaption, the method first determines if the new mesh is sufficiently unbalanced to warrant a repartitioning. If so, the adapted mesh is repartitioned, with new partitions assigned to processors so that the redistribution cost is minimized. The new partitions are accepted only if the remapping cost is compensated by the improved load balance. Results indicate that this strategy is effective for large-scale scientific computations on distributed-memory multiprocessors.

Biswas, Rupak↗

Balancing Contention and Synchronization on the Intel Paragon

The Intel Paragon is a mesh-connected distributed memory parallel computer. It uses an oblivious and deterministic message routing algorithm: this permits us to develop highly optimized schedules for frequently needed communication patterns. The complete exchange is one such pattern. Several approaches are available for carrying it out on the mesh. We study an algorithm developed by Scott. This algorithm assumes that a communication link can carry one message at a time and that a node can only transmit one message at a time. It requires global synchronization to enforce a schedule of transmissions. Unfortunately global synchronization has substantial overhead on the Paragon. At the same time the powerful interconnection mechanism of this machine permits 2 or 3 messages to share a communication link with minor overhead. It can also overlap multiple message transmission from the same node to some extent. We develop a generalization of Scott's algorithm that executes complete exchange with a prescribed contention. Schedules that incur greater contention require fewer synchronization steps. This permits us to tradeoff contention against synchronization overhead. We describe the performance of this algorithm and compare it with Scott's original algorithm as well as with a naive algorithm that does not take interconnection structure into account. The Bounded contention algorithm is always better than Scott's algorithm and outperforms the naive algorithm for all but the smallest message sizes. The naive algorithm fails to work on meshes larger than 12 x 12. These results show that due consideration of processor interconnect and machine performance parameters is necessary to obtain peak performance from the Paragon and its successor mesh machines.

Bokhari, Shahid H.↗

A Parallel Pipelined Renderer for the Time-Varying Volume Data

This paper presents a strategy for efficiently rendering time-varying volume data sets on a distributed-memory parallel computer. Time-varying volume data take large storage space and visualizing them requires reading large files continuously or periodically throughout the course of the visualization process. Instead of using all the processors to collectively render one volume at a time, a pipelined rendering process is formed by partitioning processors into groups to render multiple volumes concurrently. In this way, the overall rendering time may be greatly reduced because the pipelined rendering tasks are overlapped with the I/O required to load each volume into a group of processors; moreover, parallelization overhead may be reduced as a result of partitioning the processors. We modify an existing parallel volume renderer to exploit various levels of rendering parallelism and to study how the partitioning of processors may lead to optimal rendering performance. Two factors which are important to the overall execution time are re-source utilization efficiency and pipeline startup latency. The optimal partitioning configuration is the one that balances these two factors. Tests on Intel Paragon computers show that in general optimal partitionings do exist for a given rendering task and result in 40-50% saving in overall rendering time.

Chiueh, Tzi-Cker↗

Some Examples of the Applications of the Transonic and Supersonic Area Rules to the Prediction of Wave Drag

The experimental wave drags of bodies and wing-body combinations over a wide range of Mach numbers are compared with the computed drags utilizing a 24-term Fourier series application of the supersonic area rule and with the results of equivalent-body tests. The results indicate that the equivalent-body technique provides a good method for predicting the wave drag of certain wing-body combinations at and below a Mach number of 1. At Mach numbers greater than 1, the equivalent-body wave drags can be misleading. The wave drags computed using the supersonic area rule are shown to be in best agreement with the experimental results for configurations employing the thinnest wings. The wave drags for the bodies of revolution presented in this report are predicted to a greater degree of accuracy by using the frontal projections of oblique areas than by using normal areas. A rapid method of computing wing area distributions and area-distribution slopes is given in an appendix.

Nelson, Robert L.↗

Aerodynamic Shape Optimization Using A Combined Distributed/Shared Memory Paradigm

Current parallel computational approaches involve distributed and shared memory paradigms. In the distributed memory paradigm, each processor has its own independent memory. Message passing typically uses a function library such as MPI or PVM. In the shared memory paradigm, such as that used on the SGI Origin 2000 machine, compiler directives are used to instruct the compiler to schedule multiple threads to perform calculations. In this paradigm, it must be assured that processors (threads) do not simultaneously access regions of memory in such away that errors would occur. This paper utilizes the latest version of the SGI MPI function library to combine the two parallelization paradigms to perform aerodynamic shape optimization of a generic wing/body.

Cheung, Samson↗

A Java-Enabled Interactive Graphical Gas Turbine Propulsion System Simulator

This paper describes a gas turbine simulation system which utilizes the newly developed Java language environment software system. The system provides an interactive graphical environment which allows the quick and efficient construction and analysis of arbitrary gas turbine propulsion systems. The simulation system couples a graphical user interface, developed using the Java Abstract Window Toolkit, and a transient, space- averaged, aero-thermodynamic gas turbine analysis method, both entirely coded in the Java language. The combined package provides analytical, graphical and data management tools which allow the user to construct and control engine simulations by manipulating graphical objects on the computer display screen. Distributed simulations, including parallel processing and distributed database access across the Internet and World-Wide Web (WWW), are made possible through services provided by the Java environment.

Reed, John A.↗

Assimilation of Surface Temperature in Land Surface Models

Hydrological models have been calibrated and validated using catchment streamflows. However, using a point measurement does not guarantee correct spatial distribution of model computed heat fluxes, soil moisture and surface temperatures. With the advent of satellites in the late 70s, surface temperature is being measured two to four times a day from various satellite sensors and different platforms. The purpose of this paper is to demonstrate use of satellite surface temperature in (a) validation of model computed surface temperatures and (b) assimilation of satellite surface temperatures into a hydrological model in order to improve the prediction accuracy of soil moistures and heat fluxes. The assimilation is carried out by comparing the satellite and the model produced surface temperatures and setting the "true"temperature midway between the two values. Based on this "true" surface temperature, the physical relationships of water and energy balance are used to reset the other variables. This is a case of nudging the water and energy balance variables so that they are consistent with each other and the true" surface temperature. The potential of this assimilation scheme is demonstrated in the form of various experiments that highlight the various aspects. This study is carried over the Red-Arkansas basin in the southern United States (a 5 deg X 10 deg area) over a time period of a year (August 1987 - July 1988). The land surface hydrological model is run on an hourly time step. The results show that satellite surface temperature assimilation improves the accuracy of the computed surface soil moisture remarkably.

Lakshmi, Venkataraman↗

Evaluation of SAGE II and Balloon-Borne Stratospheric Aerosol Measurements: Evaluation of Aerosol Measurements from SAGE II, HALOE, and Balloonborne Optical Particle Counters

Stratospheric aerosol measurements from the University of Wyoming balloonborne optical particle counters (OPCs), the Stratospheric Aerosol and Gas Experiment (SAGE) II, and the Halogen Occultation Experiment (HALOE) were compared in the period 1982-2000, when measurements were available. The OPCs measure aerosol size distributions, and HALOE multiwavelength (2.45-5.26 micrometers) extinction measurements can be used to retrieve aerosol size distributions. Aerosol extinctions at the SAGE II wavelengths (0.386-1.02 micrometers) were computed from these size distributions and compared to SAGE II measurements. In addition, surface areas derived from all three experiments were compared. While the overall impression from these results is encouraging, the agreement can change with latitude, altitude, time, and parameter. In the broadest sense, these comparisons fall into two categories: high aerosol loading (volcanic periods) and low aerosol loading (background periods and altitudes above 25 km). When the aerosol amount was low, SAGE II and HALOE extinctions were higher than the OPC estimates, while the SAGE II surface areas were lower than HALOE and the OPCS. Under high loading conditions all three instruments mutually agree to within 50%.

Hervig, Mark↗