Search NASA⌕ Search

SEARCH · Search NASA

Results for “distributed parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42

Plasma rest frame distributions of suprathermal ions in the earth's foreshock region

The paper presents rest frame ion distributions computed from three-dimensional observations of upstream suprathermal ions made by the University of Iowa Quadrispherical Lepedea on ISEE-1. The observations are for a single inbound midmorning pass starting upstream from the ion foreshock and continuing across the quasi-spherical bow shock into the magnetosheath. The crossing of the ion foreshock boundary is marked by a several-minute burst of ions of temperature 100-200 eV moving along the IMF away from the bow shock at 500 km/s relative to the solar wind. The observation of these reflected ions is followed by an extended interval of diffuse ions of temperatures 2-3 keV flowing at about 250 km/s relative to the solar wind and persisting until the bow shock is crossed. Both types of suprathermal ions constitute roughly 2% of the total ion density and carry a parallel heat flux of 0.01 ergs/sq cm-s.

Sentman, D. D.↗

Application of the FUN3D Unstructured-Grid Navier-Stokes Solver to the 4th AIAA Drag Prediction Workshop Cases

FUN3D Navier-Stokes solutions were computed for the 4th AIAA Drag Prediction Workshop grid convergence study, downwash study, and Reynolds number study on a set of node-based mixed-element grids. All of the baseline tetrahedral grids were generated with the VGRID (developmental) advancing-layer and advancing-front grid generation software package following the gridding guidelines developed for the workshop. With maximum grid sizes exceeding 100 million nodes, the grid convergence study was particularly challenging for the node-based unstructured grid generators and flow solvers. At the time of the workshop, the super-fine grid with 105 million nodes and 600 million elements was the largest grid known to have been generated using VGRID. FUN3D Version 11.0 has a completely new pre- and post-processing paradigm that has been incorporated directly into the solver and functions entirely in a parallel, distributed memory environment. This feature allowed for practical pre-processing and solution times on the largest unstructured-grid size requested for the workshop. For the constant-lift grid convergence case, the convergence of total drag is approximately second-order on the finest three grids. The variation in total drag between the finest two grids is only 2 counts. At the finest grid levels, only small variations in wing and tail pressure distributions are seen with grid refinement. Similarly, a small wing side-of-body separation also shows little variation at the finest grid levels. Overall, the FUN3D results compare well with the structured-grid code CFL3D. The FUN3D downwash study and Reynolds number study results compare well with the range of results shown in the workshop presentations.

Lee-Rausch, Elizabeth M.↗

Efficient Process Migration for Parallel Processing on Non-Dedicated Networks of Workstations

This paper presents the design and preliminary implementation of MpPVM, a software system that supports process migration for PVM application programs in a non-dedicated heterogeneous computing environment. New concepts of migration point as well as migration point analysis and necessary data analysis are introduced. In MpPVM, process migrations occur only at previously inserted migration points. Migration point analysis determines appropriate locations to insert migration points; whereas, necessary data analysis provides a minimum set of variables to be transferred at each migration pint. A new methodology to perform reliable point-to-point data communications in a migration environment is also discussed. Finally, a preliminary implementation of MpPVM and its experimental results are presented, showing the correctness and promising performance of our process migration mechanism in a scalable non-dedicated heterogeneous computing environment. While MpPVM is developed on top of PVM, the process migration methodology introduced in this study is general and can be applied to any distributed software environment.

Chanchio, Kasidit↗

BUCLASP 3: A computer program for stresses and buckling of heated composite stiffened panels and other structures, user's manual

The use of the computer program BUCLASP3 is described. The code is intended for thermal stress and instability analyses of structures such as unidirectionally stiffened panels. There are two types of instability analyses that can be effected by PAINT; (1) thermal buckling, and (2) buckling due to a specified inplane biaxial loading. Any structure that has a constant cross section in one direction, that may be idealized as an assemblage of beam elements and laminated flat and curved plate strip-elements can be analyzed. The two parallel ends of the panel must be simply supported, whereas arbitrary elastic boundary conditions may be imposed along any one or both external longitudinal side. Any variation in the temperature rise (from ambient) through the cross section of a panel is considered in the analyses but it must be assumed that in the longitudinal direction the temperature field is constant. Load distributions for the externally applied inplane biaxial loads are similar in nature to the permissible temperature field.

Tripp, L. L.↗

An operating system for future aerospace vehicle computer systems

The requirements for future aerospace vehicle computer operating systems are examined in this paper. The computer architecture is assumed to be distributed with a local area network connecting the nodes. Each node is assumed to provide a specific functionality. The network provides for communication so that the overall tasks of the vehicle are accomplished. The O/S structure is based upon the concept of objects. The mechanisms for integrating node unique objects with node common objects in order to implement both the autonomy and the cooperation between nodes is developed. The requirements for time critical performance and reliability and recovery are discussed. Time critical performance impacts all parts of the distributed operating system; e.g., its structure, the functional design of its objects, the language structure, etc. Throughout the paper the tradeoffs - concurrency, language structure, object recovery, binding, file structure, communication protocol, programmer freedom, etc. - are considered to arrive at a feasible, maximum performance design. Reliability of the network system is considered. A parallel multipath bus structure is proposed for the control of delivery time for time critical messages. The architecture also supports immediate recovery for the time critical message system after a communication failure.

Foudriat, E. C.↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

BLASIM - A computational tool to assess ice impact damage on engine blades

A portable computer code called BLASIM is developed at NASA LeRC to assess the ice impact damage on aircraft engine blades. In addition to the ice impact analyses, the code is also capable of carrying out static, dynamic, resonance margin and flutter analyses. The blade can be solid, hollow, superhybrid or composite material. An optional preprocessor (input generator) is also developed to generate input to the code through interactive process. The blade geometry can be defined either by a series of airfoils at discrete input stations or by a finite element grid. The code employs a coarse fixed finite element mesh with triangular plate finite elements and has quick turnaround time. The ice piece is modeled as an equivalent spherical object and has the velocity opposite to that of the aircraft with direction parallel to the engine axis. For the local impact damage assessment, the impact force is considered as a distributed load acting over a region around the impact point and the average radial strain of the finite elements along the leading edge is taken as a measure of the local damage. To estimate the damage at the blade root, the impact is considered to be an impulse and a combined stress failure criteria is employed. Parametric studies for local and root ice impact damage, and post-impact dynamics are discussed for solid and composite blades.

Reddy, E. S.↗

Unified Models of Turbulence and Nonlinear Wave Evolution in the Extended Solar Corona and Solar Wind

The PI (Cranmer) and Co-I (A. van Ballegooijen) made significant progress toward the goal of building a "unified model" of the dominant physical processes responsible for the acceleration of the solar wind. The approach outlined in the original proposal comprised two complementary pieces: (1) to further investigate individual physical processes under realistic coronal and solar wind conditions, and (2) to extract the dominant physical effects from simulations and apply them to a one-dimensional and time-independent model of plasma heating and acceleration. The accomplishments in the report period are thus divided into these two categories: 1a. Focused Study of Kinetic MHD Turbulence. We have developed a model of magnetohydrodynamic (MHD) turbulence in the extended solar corona that contains the effects of collisionless dissipation and anisotropic particle heating. A turbulent cascade is one possible way of generating small-scale fluctuations (easy to dissipate/heat) from a pre-existing population of low-frequency Alfven waves (difficult to dissipate/heat). We modeled the cascade as a combination of advection and diffusion in wavenumber space. The dominant spectral transfer occurs in the direction perpendicular to the background magnetic field. As expected from earlier models, this leads to a highly anisotropic fluctuation spectrum with a rapidly decaying tail in the parallel wavenumber direction. The wave power that decays to high enough frequencies to become ion cyclotron resonant depends on the relative strengths of advection and diffusion in the cascade. For the most realistic values of these parameters, though, there is insufficient power to heat protons and heavy ions. The dominant oblique waves undergo Landau damping, which implies strong parallel electron heating. We thus investigated the nonlinear evolution of the electron velocity distributions (VDFs) into parallel beams and discrete phase-space holes (similar to those seen in the terrestrial magnetosphere) which are an alternate means of heating protons via stochastic interactions similar to particle-particle collisions. 1b. Focused Study of the Multi-Mode Detailed Balance Formalism. The PI began to explore the feasibility of using the "weak turbulence," or detailed-balance theory of Tsytovich, Melrose, and others to encompass the relevant physics of the solar wind. This study did not go far, however, because if the "strong" MHD turbulence discussed above is a dominant player in the wind's acceleration region, this formalism is inherently not applicable to the corona. We will continue to study the various published approaches to the weak turbulence formalism, especially with an eye on ways to parameterize nonlinear wave reflection rates. 2. Building the Unified Model Code Architecture. We have begun developing the computational model of a time-steady open flux tube in the extended corona. The model will be "unified" in the sense that it will include (simultaneously for the first time) as many of the various proposed physical processes as possible, all on equal footing. To retain this generality, we have formulated the problem in two interconnected parts: a completely kinetic model for the particles, using the Monte Carlo approach, and a finite-difference approach for the self-consistent fluctuation spectra. The two codes are run sequentially and iteratively until complete consistency is achieved. The current version of the Monte Carlo code incorporates gravity, the zero-current electric field, magnetic mirroring, and collisions. The fluctuation code incorporates WKJ3 wave action conservation and the cascade/dissipation processes discussed above. The codes are being run for various test problems with known solutions. Planned additions to the codes include prescriptions for nonlinear wave steepening, kinetic velocity-space diffusion, and multi-mode coupling (including reflection and refraction).

Cranmer, Steven R.↗

Communication Studies of DMP and SMP Machines

Understanding the interplay between machines and problems is key to obtaining high performance on parallel machines. This paper investigates the interplay between programming paradigms and communication capabilities of parallel machines. In particular, we explicate the communication capabilities of the IBM SP-2 distributed-memory multiprocessor and the SGI PowerCHALLENGEarray symmetric multiprocessor. Two benchmark problems of bitonic sorting and Fast Fourier Transform are selected for experiments. Communication-efficient algorithms are developed to exploit the overlapping capabilities of the machines. Programs are written in Message-Passing Interface for portability and identical codes are used for both machines. Various data sizes and message sizes are used to test the machines' communication capabilities. Experimental results indicate that the communication performance of the multiprocessors are consistent with the size of messages. The SP-2 is sensitive to message size but yields a much higher communication overlapping because of the communication co-processor. The PowerCHALLENGEarray is not highly sensitive to message size and yields a low communication overlapping. Bitonic sorting yields lower performance compared to FFT due to a smaller computation-to-communication ratio.

Sohn, Andrew↗

Two-Dimensional Sequential and Concurrent Finite Element Analysis of Unstiffened and Stiffened Aluminum and Composite Panels with Hole

The results of a detailed investigation of the distribution of stresses in aluminum and composite panels subjected to uniform end shortening are presented. The focus problem is a rectangular panel with two longitudinal stiffeners, and an inner stiffener discontinuous at a central hole in the panel. The influence of the stiffeners on the stresses is evaluated through a two-dimensional global finite element analysis in the absence or presence of the hole. Contrary to the physical feel, it is found that the maximum stresses from the glocal analysis for both stiffened aluminum and composite panels are greater than the corresponding stresses for the unstiffened panels. The inner discontinuous stiffener causes a greater increase in stresses than the reduction provided by the two outer stiffeners. A detailed layer-by-layer study of stresses around the hole is also presented for both unstiffened and stiffened composite panels. A parallel equation solver is used for the global system of equations since the computational time is far less than that using a sequential scheme. A parallel Choleski method with up to 16 processors is used on Flex/32 Multicomputer at NASA Langley Research Center. The parallel computing results are summarized and include the computational times, speedups, bandwidths, and their inter-relationships for the panel problems. It is found that the computational time for the Choleski method decreases with a decrease in bandwidth, and better speedups result as the bandwidth increases.

Razzaq, Zia↗

Computational Modeling of Graphite Degradation due to Molten Salt Infiltration and Wear

Molten-salt reactors (MSRs) represent a promising next-generation reactor design, with graphite serving as a moderator and/or reflector in several designs. However, due to limited experimental data and operational experience, a technical understanding of the structural integrity of graphite in molten salt environments remains incomplete. This report presents a modeling-based evaluation of graphite degradation in MSR environments, focusing on the effects of salt infiltration in fuel salt-based designs and surface wear in pebble bed reactor designs. The objective of this study is to enhance understanding of the structural integrity challenges posed by these degradation mechanisms and to provide a framework for assessing graphite behavior in MSRs. The first part of the report investigates the phenomenon of molten salt infiltration into graphite. This infiltration occurs when molten salt permeates the interconnected pore structure of the graphite moderator, driven by factors such as pressure differentials and the physical properties of both the salt and graphite. The infiltration process is influenced by characteristics of the pore structure, viscosity of the molten salt, and the interfacial energies between the graphite, salt, and the atmosphere within the graphite pore. Utilizing a coupled multiphysics modeling approach with Grizzly software, the study evaluates the stress induced by internal heat sources due to infiltration, which can lead to structural concerns. This evaluation is crucial for understanding how infiltration affects the mechanical integrity of graphite components in MSRs. The study considers the Molten-Salt Reactor Experiment (MSRE) graphite stringer geometry due to the availability of relevant data. Through detailed finite element analysis, the study examines stress distributions at varying infiltration percentages, revealing that stress levels increase with higher amounts of infiltration. Rare-event simulations, using the parallel subset simulation (PSS) framework, further quantify the failure probabilities under input uncertainties, with a user-specified failure metric. The PSS framework also identifies critical input parameters that significantly affect the stress values, including infiltration amount, thermal conductivity, and power density. Additionally, considering realistic reactor scenarios, the analysis was performed to account for the combined effects of radiation and infiltration, and modeling strategies on how to analyze new reactor designs or new graphite grades are discussed. The second part of the report focuses on wear mechanisms in pebble bed-based MSRs. As graphite fuel pebbles interact with the graphite reflector block, wear can result in material loss and the formation of surface defects, which may act as stress concentrators. A similar multiphysics modeling framework is employed to assess the impact of wear on the structural integrity of graphite components. This study considers a generic fluoride-cooled high-temperature reactor (gFHR) design due to the availability of comprehensive data. Worst-case scenario dimensions of the reflector blocks were analyzed under thermal and radiation conditions. Subsequently, wear in the form of idealized pits and grooves is modeled on the inner surface of the graphite block, with the maximum stress from previous simulations. The simulations show that groove-type defects are more detrimental than pits, leading to higher stress concentrations. Considering worst-case simulation scenarios and experimental wear rates, it was determined that the formation of a surface defect critical enough to affect the stress may not be possible in a gFHR design. Overall, the findings of this research contribute to the development of robust modeling tools for predicting graphite behavior under various operational conditions in MSRs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Method and Apparatus of Multiplexing and Acquiring Data from Multiple Optical Fibers Using a Single Data Channel of an Optical Frequency-Domain Reflectometry (OFDR) System

A method and system for multiplexing a network of parallel fiber Bragg grating (FBG) sensor-fibers to a single acquisition channel of a closed Michelson interferometer system via a fiber splitter by distinguishing each branch of fiber sensors in the spatial domain. On each branch of the splitter, the fibers have a specific pre-determined length, effectively separating each branch of fiber sensors spatially. In the spatial domain the fiber branches are seen as part of one acquisition channel on the interrogation system. However, the FBG-reference arm beat frequency information for each fiber is retained. Since the beat frequency is generated between the reference arm, the effective fiber length of each successive branch includes the entire length of the preceding branch. The multiple branches are seen as one fiber having three segments where the segments can be resolved. This greatly simplifies optical, electronic and computational complexity, and is especially suited for use in multiplexed or branched OFS networks for SHM of large and/or distributed structures which need a lot of measurement points.

Parker, Jr., Allen R↗

Transition Analysis for the CRM-NLF Wind Tunnel Configuration

This paper presents the results of an ongoing study into the linear stability characteristics of the boundary layer flow over the common research model with natural laminar flow (CRMNLF) aircraft configuration. The flow conditions match selected test conditions from a recent wind tunnel experiment in the National Transonic Facility at the NASA Langley Research Center. Previous work involving parallel stability computations of a boundary layer flow based on the conical flow approximation has shown that the measured onset of laminar-turbulent transition during the experiments can be correlated with the linear amplification of Tollmien- Schlichting (TS) and stationary crossflow (CF) instabilities in the swept wing boundary layer. Here, we examine the effects of the simplifying approximations in both basic state computation and the stability analysis, with the goal of quantifying the resulting changes in the N-factor correlations. Specifically, the basic states are computed by using full Navier-Stokes equations and the stability analysis is performed by using a nonorthogonal coordinate system that allows a clear distinction between planar TS and CF instabilities. Furthermore, the effects of curvature and nonparallel mean flow have been included in the stability computations based on the parabolized stability equations (PSE). The fully turbulent Reynolds-Averaged-Navier-Stokes (RANS) mean flow solutions show good agreement with the measured wall pressure distribution. Viscous-inviscid interactive effects are observed to be important because the shock fronts along the suction surface are influenced by the imposed transition front. The stability results confirm the previous findings related to TS amplification within the inboard region of the wing and the dominance of stationary CF modes in the outboard region. However, given the close proximity of the measured transition front and the dual shock system within the outer part of the wing, the onset of transition may well be shock limited within the outboard region. In general, the transition criterion based on the dual N-factor method with N TS = N CF = 6 is reasonably successful at correlating with the measured transition fronts at Re MAC = 15 million and AoA = 1.5, 2 degrees; however, the low values of the correlating N-factors at Re MAC = 17.5 million support the hypothesis that the measured transition at the higher Reynolds number may have been strongly influenced by the merging of turbulent wedges that originate from surface imperfections near the leading edge.

Boundary layer transition↗

Transition Analysis for the CRM-NLF Wind Tunnel Configuration

This paper presents the results of an ongoing study into the linear stability characteristics of the boundary layer flow over the common research model with natural laminar flow (CRMNLF) aircraft configuration. The flow conditions match selected test conditions from a recent wind tunnel experiment in the National Transonic Facility at the NASA Langley Research Center. Previous work involving parallel stability computations of a boundary layer flow based on the conical flow approximation has shown that the measured onset of laminar-turbulent transition during the experiments can be correlated with the linear amplification of Tollmien- Schlichting (TS) and stationary crossflow (CF) instabilities in the swept wing boundary layer. Here, we examine the effects of the simplifying approximations in both basic state computation and the stability analysis, with the goal of quantifying the resulting changes in the N-factor correlations. Specifically, the basic states are computed by using full Navier-Stokes equations and the stability analysis is performed by using a nonorthogonal coordinate system that allows a clear distinction between planar TS and CF instabilities. Furthermore, the effects of curvature and nonparallel mean flow have been included in the stability computations based on the parabolized stability equations (PSE). The fully turbulent Reynolds-Averaged-Navier-Stokes (RANS) mean flow solutions show good agreement with the measured wall pressure distribution. Viscous-inviscid interactive effects are observed to be important because the shock fronts along the suction surface are influenced by the imposed transition front. The stability results confirm the previous findings related to TS amplification within the inboard region of the wing and the dominance of stationary CF modes in the outboard region. However, given the close proximity of the measured transition front and the dual shock system within the outer part of the wing, the onset of transition may well be shock limited within the outboard region. In general, the transition criterion based on the dual N-factor method with N TS = N CF = 6 is reasonably successful at correlating with the measured transition fronts at Re MAC = 15 million and AoA = 1.5, 2 degrees; however, the low values of the correlating N-factors at Re MAC = 17.5 million support the hypothesis that the measured transition at the higher Reynolds number may have been strongly influenced by the merging of turbulent wedges that originate from surface imperfections near the leading edge.

Boundary layer transition↗

New description of charged particle propagation in random magnetic fields

When charged particles spiral along a large constant magnetic field, their trajectories are scattered by random components that are superposed on the guiding field. In the simplest analysis of this situation, scattering causes the particles to diffuse parallel to the guiding field. At the next level of approximation, moving pulses that correspond to a coherent mode of propagation are present, but they are represented by delta-functions whose infinitely narrow width makes no sense physically and is inconsistent with the finite duration of coherent pulses observed in solar energetic particle events. To derive a more realistic description, the transport problem is formulated in terms of 4 x 4 matrices, which derive from a representation of the particle distribution function in terms of eigenfunctions of the scattering operator, and which lead to useful approximations that give explicit predictions of the detailed evolution not only of the coherent pulses, but also of the diffusive wake. More specifically, the new description embodies a simple convolution of a narrow Gaussian with the solutions above that involve delta-functions, but with a slightly reduced coherent velocity. The validity of these approximations, which can easily be calculated on a desktop computer, has been exhaustively confirmed by comparison with results of Monte Carlo simulations which kept track of 50 million particles and which were carried out on the Maspar computer at Goddard Space Flight Center.

Earl, James A.↗

The characteristic of the magnetopause reconnection X-line deduced from low-altitude satellite observations of cusp ions

We present an analysis of a 'quasi-steady' cusp ion dispersion signature observed at low altitudes. We reconstruct the field-parallel part of the Cowley-D ion distribution function, injected into the open low-latitude boundary layer (LLBL) in the vicinity of the reconnection X-line. From this we find the field parallel magnetosheath flow at the X-line was only 20 +/- 60 km/s, placing the reconnection site close to the flow streamline which is perpendicular to the magnetosheath field. Using interplanetary data and assuming the subsolar magnetopause is in pressure balance, we derive a wealth of information about the X-line, including: the density, flow, magnetic field and Alfven speed of the magnetosheath; the magnetic shear across the X-line; the de-Hoffman Teller speed with which field lines emerge from the X-line; the magnetospheric field; and the ion transmission factor across the magnetopause. The results indicate that some heating takes place near the X-line as the ions cross the magnetopause, and that sheath densities may be reduced in a plasma depletion layer. We also compute the reconnection rate. Despite its quasi-steady appearance on an ion spectrogram, this cusp is found to reveal a large pulse of enhanced reconnection rate.

Lockwood, M.↗

Volumetric 3D Display System with Static Screen

Current display technology has relied on flat, 2D screens that cannot truly convey the third dimension of visual information: depth. In contrast to conventional visualization that is primarily based on 2D flat screens, the volumetric 3D display possesses a true 3D display volume, and places physically each 3D voxel in displayed 3D images at the true 3D (x,y,z) spatial position. Each voxel, analogous to a pixel in a 2D image, emits light from that position to form a real 3D image in the eyes of the viewers. Such true volumetric 3D display technology provides both physiological (accommodation, convergence, binocular disparity, and motion parallax) and psychological (image size, linear perspective, shading, brightness, etc.) depth cues to human visual systems to help in the perception of 3D objects. In a volumetric 3D display, viewers can watch the displayed 3D images from a completely 360 view without using any special eyewear. The volumetric 3D display techniques may lead to a quantum leap in information display technology and can dramatically change the ways humans interact with computers, which can lead to significant improvements in the efficiency of learning and knowledge management processes. Within a block of glass, a large amount of tiny dots of voxels are created by using a recently available machining technique called laser subsurface engraving (LSE). The LSE is able to produce tiny physical crack points (as small as 0.05 mm in diameter) at any (x,y,z) location within the cube of transparent material. The crack dots, when illuminated by a light source, scatter the light around and form visible voxels within the 3D volume. The locations of these tiny voxels are strategically determined such that each can be illuminated by a light ray from a high-resolution digital mirror device (DMD) light engine. The distribution of these voxels occupies the full display volume within the static 3D glass screen. This design eliminates any moving screen seen in previous approaches, so there is no image jitter, and has an inherent parallel mechanism for 3D voxel addressing. High spatial resolution is possible with a full color display being easy to implement. The system is low-cost and low-maintenance.

Geng, Jason↗

Integrating Cache Performance Modeling and Tuning Support in Parallelization Tools

With the resurgence of distributed shared memory (DSM) systems based on cache-coherent Non Uniform Memory Access (ccNUMA) architectures and increasing disparity between memory and processors speeds, data locality overheads are becoming the greatest bottlenecks in the way of realizing potential high performance of these systems. While parallelization tools and compilers facilitate the users in porting their sequential applications to a DSM system, a lot of time and effort is needed to tune the memory performance of these applications to achieve reasonable speedup. In this paper, we show that integrating cache performance modeling and tuning support within a parallelization environment can alleviate this problem. The Cache Performance Modeling and Prediction Tool (CPMP), employs trace-driven simulation techniques without the overhead of generating and managing detailed address traces. CPMP predicts the cache performance impact of source code level "what-if" modifications in a program to assist a user in the tuning process. CPMP is built on top of a customized version of the Computer Aided Parallelization Tools (CAPTools) environment. Finally, we demonstrate how CPMP can be applied to tune a real Computational Fluid Dynamics (CFD) application.

Waheed, Abdul↗