Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62

Transition of a Three-Dimensional Unsteady Viscous Flow Analysis from a Research Environment to the Design Environment

The advent of advanced computer architectures and parallel computing have led to a revolutionary change in the design process for turbomachinery components. Two- and three-dimensional steady-state computational flow procedures are now routinely used in the early stages of design. Unsteady flow analyses, however, are just beginning to be incorporated into design systems. This paper outlines the transition of a three-dimensional unsteady viscous flow analysis from the research environment into the design environment. The test case used to demonstrate the analysis is the full turbine system (high-pressure turbine, inter-turbine duct and low-pressure turbine) from an advanced turboprop engine.

Dorney, Suzanne↗

Reengineering the Project Design Process

In response to NASA's goal of working faster, better and cheaper, JPL has developed extensive plans to minimize cost, maximize customer and employee satisfaction, and implement small- and moderate-size missions. These plans include improved management structures and processes, enhanced technical design processes, the incorporation of new technology, and the development of more economical space- and ground-system designs. The Laboratory's new Flight Projects Implementation Office has been chartered to oversee these innovations and the reengineering of JPL's project design process, including establishment of the Project Design Center and the Flight System Testbed. Reengineering at JPL implies a cultural change whereby the character of its design process will change from sequential to concurrent and from hierarchical to parallel. The Project Design Center will support missions offering high science return, design to cost, demonstrations of new technology, and rapid development. Its computer-supported environment will foster high-fidelity project life-cycle development and cost estimating.

Jet Propulsion Laboratory JPL project design concu↗

Photometer for tracking a moving light source

A photometer that tracks a path of a moving light source with little or no motion of the photometer components. The system includes a non-moving, truncated paraboloid of revolution, having a paraboloid axis, a paraboloid axis, a small entrance aperture, a larger exit aperture and a light-reflecting inner surface, that receives and reflects light in a direction substantially parallel to the paraboloid axis. The system also includes a light processing filter to receive and process the redirected light, and to issue the processed, redirected light as processed light, and an array of light receiving elements, at least one of which receives and measures an associated intensity of a portion of the processed light. The system tracks a light source moving along a path and produces a corresponding curvilinear image of the light source path on the array of light receiving elements. Undesired light wavelengths from the light source may be removed by coating a selected portion of the reflecting inner surface or another light receiving surface with a coating that absorbs incident light in the undesired wavelength range.

Strawa, Anthony W.↗

Cartesian Off-Body Grid Adaption for Viscous Time- Accurate Flow Simulation

An improved solution adaption capability has been implemented in the OVERFLOW overset grid CFD code. Building on the Cartesian off-body approach inherent in OVERFLOW and the original adaptive refinement method developed by Meakin, the new scheme provides for automated creation of multiple levels of finer Cartesian grids. Refinement can be based on the undivided second-difference of the flow solution variables, or on a specific flow quantity such as vorticity. Coupled with load-balancing and an inmemory solution interpolation procedure, the adaption process provides very good performance for time-accurate simulations on parallel compute platforms. A method of using refined, thin body-fitted grids combined with adaption in the off-body grids is presented, which maximizes the part of the domain subject to adaption. Two- and three-dimensional examples are used to illustrate the effectiveness and performance of the adaption scheme.

Buning, Pieter G.↗

INSPiRE – An Approach to Mission Quality Management using Network Slicing for Space Applications

Managing traffic between the Earth-Moon and Earth-Mars is a complex process requiring significant investment in resources and expertise at NASA. INSPiRE improves the performance of space networks by enabling a dynamic re-configuration process that works for any mixed topology over a heterogeneous and multi-vendor network. To achieve the desired functionality, INSPiRE incorporates a set of algorithms, machine learning processes, and policy inference to handle unpredictable, disruptive events. INSPiRE draws parallels from the current notion of the 3GPP (5G and beyond) Network Slicing approach, where the same physical network divides into several virtual networks, and for each of these virtual networks, there is a guaranteed Quality of Service for the missions that they serve.

cognitive communications↗

Computational Study of Oxidative Etch Pitting in FiberForm and the Effect on Its Material Properties

Erosion of carbon due to oxidation does not occur uniformly but through the formation of localized etch pits because of active surface sites. These active sites are formed due to the presence of atomic defects on the carbon surface, and have much higher reactivity compared to average non-defective sites. Thus, these active sites are first to react during ablation, resulting in their removal. This causes all the neighboring atoms to be defective and increase their reactivity, thus leading to the localized carbon removal around these “active” sites. In this manner, these highly reactive defective sites serve as nucleation sites for the formation and growth of etch pits with potentially detrimental effects on the structural integrity. In order to understand the influence of these etch pits on the material properties of carbon fiber microstructures, we have developed a new capability within direct simulation Monte Carlo (DSMC) to capture the etch pit formation process. This capability is developed within the DSMC code SPARTA (Stochastic PArallel Rarefied-gas Time-accurate Analyzer) and can model the material removal in presence of active sites leading to the formation of etch pits. The focus of the current work will be to study the effect of etch pits on the material properties of FiberForm, a commonly used base material within many thermal protection system materials (TPS). The microstructure of virgin FiberForm obtained directly from X-ray microtomography experiments is used within SPARTA to obtain the ablated geometries with etch pits. These pitted microstructures are then imported within the Porous Microstructure Analysis (PuMA) software and various material properties such as elasticity, thermal conductivity, and permeability are computed. The variation of these properties as a result of the complex evolution of the surface topology due to etch pit formation is studied and analyzed. Furthermore, the effect of pitting is compared to the case of shrinking fibers, which has been the standard for modelling ablation of carbon structures; and significant differences are observed. Thus, such a physically realistic modeling of material removal through the formation of etch pits will be helpful in predicting the degradation of carbon-based TPS more accurately during oxidation; as well as other mechanisms such as spallation, which involves the removal of chunks of material into the flow due to etch pit growth. This will ultimately improve our understanding of the failure modes in these materials due to ablation.

Carbon Ablators↗

Computational Study of Oxidative Etch Pitting in FiberForm and the Effect on Its Material Properties

Erosion of carbon due to oxidation does not occur uniformly but through the formation of localized etch pits because of active surface sites. These active sites are formed due to the presence of atomic defects on the carbon surface, and have much higher reactivity compared to average non-defective sites. Thus, these active sites are first to react during ablation, resulting in their removal. This causes all the neighboring atoms to be defective and increase their reactivity, thus leading to the localized carbon removal around these “active” sites. In this manner, these highly reactive defective sites serve as nucleation sites for the formation and growth of etch pits with potentially detrimental effects on the structural integrity. In order to understand the influence of these etch pits on the material properties of carbon fiber microstructures, we have developed a new capability within direct simulation Monte Carlo (DSMC) to capture the etch pit formation process. This capability is developed within the DSMC code SPARTA (Stochastic PArallel Rarefied-gas Time-accurate Analyzer) and can model the material removal in presence of active sites leading to the formation of etch pits. The focus of the current work will be to study the effect of etch pits on the material properties of FiberForm, a commonly used base material within many thermal protection system materials (TPS). The microstructure of virgin FiberForm obtained directly from X-ray microtomography experiments is used within SPARTA to obtain the ablated geometries with etch pits. These pitted microstructures are then imported within the Porous Microstructure Analysis (PuMA) software and various material properties such as elasticity, thermal conductivity, and permeability are computed. The variation of these properties as a result of the complex evolution of the surface topology due to etch pit formation is studied and analyzed. Furthermore, the effect of pitting is compared to the case of shrinking fibers, which has been the standard for modelling ablation of carbon structures; and significant differences are observed. Thus, such a physically realistic modeling of material removal through the formation of etch pits will be helpful in predicting the degradation of carbon-based TPS more accurately during oxidation; as well as other mechanisms such as spallation, which involves the removal of chunks of material into the flow due to etch pit growth. This will ultimately improve our understanding of the failure modes in these materials due to ablation.

Carbon Ablators↗

Simulation of Oxidative Etch Pit Formation and Growth on FiberForm

Erosion of carbon surfaces due to oxidation does not occur uniformly but through the formation of localized etch pits because of active surface sites. These active sites are formed due to the presence of atomic defects on the carbon surface and have much higher reactivity compared to average non-defective sites. Thus, these active sites are the first to react during ablation, resulting in their removal. This causes all the neighboring atoms to be defective, and increasing their reactivity, thus leading to localized carbon removal around these “active” sites. In this manner, these highly reactive defects serve as nucleation sites for the formation and growth of etch pits, with detrimental effects on the structural integrity. In order to understand the influence of these etch pits on the material properties of carbon fiber microstructures, we have developed a new capability within direct simulation Monte Carlo (DSMC) to capture the etch pit formation process. This capability is developed within the DSMC code SPARTA (Stochastic PArallel Rarefied-gas Time-accurate Analyzer) and can model the material removal in the presence of active sites leading to the formation of etch pits. The focus of the current work will be to study the effect of etch pits on the material properties of FiberForm, the precursor substrate of the PICA Thermal Protection System (TPS) material. The microstructure of virgin FiberForm obtained directly from X-ray microtomography scans is used within SPARTA to obtain the ablated geometries with etch pits. These pitted microstructures are then imported into the Porous Microstructure Analysis (PuMA) software and various material properties such as thermal conductivity, elasticity, and permeability are computed. The variation of these properties because of the complex evolution of the surface topology due to the formation of etch pits is studied and analyzed. Furthermore, the effect of pitting is compared to the case of uniform radial shrinking of fibers, which has been the standard for modelling ablation of carbon structures, and significant differences are observed. Thus, a physically realistic model of material removal through the formation of etch pits will be helpful in predicting the degradation of carbon-based TPS more accurately during oxidation; as well as other mechanisms such as spallation, which involves the removal of chunks of material into the flow due to the growth of etch pits. This will ultimately improve our understanding of the failure modes in these materials due to ablation.

DSMC↗

Simulation of Oxidative Etch Pit Formation and Growth on FiberForm

Erosion of carbon surfaces due to oxidation does not occur uniformly but through the formation of localized etch pits because of active surface sites. These active sites are formed due to the presence of atomic defects on the carbon surface and have much higher reactivity compared to average non-defective sites. Thus, these active sites are the first to react during ablation, resulting in their removal. This causes all the neighboring atoms to be defective, and increasing their reactivity, thus leading to localized carbon removal around these “active” sites. In this manner, these highly reactive defects serve as nucleation sites for the formation and growth of etch pits, with detrimental effects on the structural integrity. In order to understand the influence of these etch pits on the material properties of carbon fiber microstructures, we have developed a new capability within direct simulation Monte Carlo (DSMC) to capture the etch pit formation process. This capability is developed within the DSMC code SPARTA (Stochastic PArallel Rarefied-gas Time-accurate Analyzer) and can model the material removal in the presence of active sites leading to the formation of etch pits. The focus of the current work will be to study the effect of etch pits on the material properties of FiberForm, the precursor substrate of the PICA Thermal Protection System (TPS) material. The microstructure of virgin FiberForm obtained directly from X-ray microtomography scans is used within SPARTA to obtain the ablated geometries with etch pits. These pitted microstructures are then imported into the Porous Microstructure Analysis (PuMA) software and various material properties such as thermal conductivity, elasticity, and permeability are computed. The variation of these properties because of the complex evolution of the surface topology due to the formation of etch pits is studied and analyzed. Furthermore, the effect of pitting is compared to the case of uniform radial shrinking of fibers, which has been the standard for modelling ablation of carbon structures, and significant differences are observed. Thus, a physically realistic model of material removal through the formation of etch pits will be helpful in predicting the degradation of carbon-based TPS more accurately during oxidation; as well as other mechanisms such as spallation, which involves the removal of chunks of material into the flow due to the growth of etch pits. This will ultimately improve our understanding of the failure modes in these materials due to ablation.

DSMC↗

High Performance Access to Archival Data Stored in HDF4 and HDF5 on Cloud Object Stores Without Reformatting the Files

Cloud computing offers numerous advantages for users of extensive Earth science data collections. These benefits encompass direct online access to data files and granules from any location, scalable access supporting parallel computing workflows, and flexible computing tools enabling innovative experimentation with processing techniques. However, older archival file formats designed for distinct computing systems hinder efficient access to decade-long time-series data when compared to data stored in modern cloud-optimized formats like Web Object Stores (WOS), exemplified by Amazon Web Services’ Simple Storage Service (S3). We describe DMR++ (Dataset Metadata Response plus plus), a technology facilitating efficient access to HDF5 (Hierarchical Data Format, version 5) and HDF4 files stored on WOS systems without requiring data reformatting. DMR++ achieves performance comparable to technologies like Zarr while preserving the original file structure, a substantial benefit considering the vast quantity of archival files held by organizations such as NASA. Moreover, DMR++ typically outperforms cloud-optimized versions of HDF5. Essentially an XML (Extensible Markup Language) document usually stored alongside the described data, DMR++ can also be generated on-the-fly but is generally created during data staging to the WOS. Archival files that use HDF4/5 often store large arrays of numerical data. The data in these files is often compressed, typically reducing their size by a factor of four or more. To achieve efficient access to portions of those arrays, they are 'chunked' into smaller sub-arrays, each individually compressed. The chunk size is a compromise, where spinning disks can efficiently access data in smaller chunks while S3 favors larger chunks. A simple optimization of aggregating smaller chunks that are stored adjacently, transferring them in a single access and then individually decompressing them will improve performance. NASA data pose an additional challenge: special Application Programmer Interface (API) libraries are often needed to compute some variables. These libraries are incompatible with WOS environments. Our solution involves storing computed values in the DMR++ document or a companion file, making them accessible like other variables and eliminating the need for specialized APIs. We outline specific optimizations for both satellite grid and swath data stored in HDF4-EOS2 (Earth Observing System).

James Gallagher↗

Multiprocessing the Sieve of Eratosthenes

The Sieve of Eratosthenes for finding prime numbers in recent years has seen much use as a benchmark algorithm for serial computers while its intrinsically parallel nature has gone largely unnoticed. The implementation of a parallel version of this algorithm for a real parallel computer, the Flex/32, is described and its performance discussed. It is shown that the algorithm is sensitive to several fundamental performance parameters of parallel machines, such as spawning time, signaling time, memory access, and overhead of process switching. Because of the nature of the algorithm, it is impossible to get any speedup beyond 4 or 5 processors unless some form of dynamic load balancing is employed. We describe the performance of our algorithm with and without load balancing and compare it with theoretical lower bounds and simulated results. It is straightforward to understand this algorithm and to check the final results. However, its efficient implementation on a real parallel machine requires thoughtful design, especially if dynamic load balancing is desired. The fundamental operations required by the algorithm are very simple: this means that the slightest overhead appears prominently in performance data. The Sieve thus serves not only as a very severe test of the capabilities of a parallel processor but is also an interesting challenge for the programmer.

Bokhari, S.↗

Investigating the Simulink Auto-Coding Process

Model based program design is the most clear and direct way to develop algorithms and programs for interfacing with hardware. While coding "by hand" results in a more tailored product, the ever-growing size and complexity of modern-day applications can cause the project work load to quickly become unreasonable for one programmer. This has generally been addressed by splitting the product into separate modules to allow multiple developers to work in parallel on the same project, however this introduces new potentials for errors in the process. The fluidity, reliability and robustness of the code relies on the abilities of the programmers to communicate their methods to one another; furthermore, multiple programmers invites multiple potentially differing coding styles into the same product, which can cause a loss of readability or even module incompatibility. Fortunately, Mathworks has implemented an auto-coding feature that allows programmers to design their algorithms through the use of models and diagrams in the graphical programming environment Simulink, allowing the designer to visually determine what the hardware is to do. From here, the auto-coding feature handles converting the project into another programming language. This type of approach allows the designer to clearly see how the software will be directing the hardware without the need to try and interpret large amounts of code. In addition, it speeds up the programming process, minimizing the amount of man-hours spent on a single project, thus reducing the chance of human error as well as project turnover time. One such project that has benefited from the auto-coding procedure is Ramses, a portion of the GNC flight software on-board Orion that has been implemented primarily in Simulink. Currently, however, auto-coding Ramses into C++ requires 5 hours of code generation time. This causes issues if the tool ever needs to be debugged, as this code generation will need to occur with each edit to any part of the program; additionally, this is lost time that could be spent testing and analyzing the code. This is one of the more prominent issues with the auto-coding process, and while much information is available with regard to optimizing Simulink designs to produce efficient and reliable C++ code, not much research has been made public on how to reduce the code generation time. It is of interest to develop some insight as to what causes code generation times to be so significant, and determine if there are architecture guidelines or a desirable auto-coding configuration set to assist in streamlining this step of the design process for particular applications. To address the issue at hand, the Simulink coder was studied at a foundational level. For each different component type made available by the software, the features, auto-code generation time, and the format of the generated code were analyzed and documented. Tools were developed and documented to expedite these studies, particularly in the area of automating sequential builds to ensure accurate data was obtained. Next, the Ramses model was examined in an attempt to determine the composition and the types of technologies used in the model. This enabled the development of a model that uses similar technologies, but takes a fraction of the time to auto-code to reduce the turnaround time for experimentation. Lastly, the model was used to run a wide array of experiments and collect data to obtain knowledge about where to search for bottlenecks in the Ramses model. The resulting contributions of the overall effort consist of an experimental model for further investigation into the subject, as well as several automation tools to assist in analyzing the model, and a reference document offering insight to the auto-coding process, including documentation of the tools used in the model analysis, data illustrating some potential problem areas in the auto-coding process, and recommendations on areas or practices in the current Ramses model that should be further investigated. Several skills were required to be built up over the course of the internship project. First and foremost, my Simulink skills have improved drastically, as much of my experience had been modeling electronic circuits as opposed to software models. Furthermore, I am now comfortable working with the Simulink Auto-coder, a tool I had never used until this summer; this tool also tested my critical thinking and C++ knowledge as I had to interpret the C++ code it was generating and attempt to understand how the Simulink model affected the generated code. I had come into the internship with a solid understanding of Matlab code, but had done very little in using it to automate tasks, particularly Simulink tasks; along the same lines, I had rarely used shell script to automate and interface with programs, which I gained a fair amount of experience with this summer, including how to use regular expression. Lastly, soft-skills are an area everyone can continuously improve on; having never worked with NASA engineers, which to me seem to be a completely different breed than what I am used to (commercial electronic engineers), I learned to utilize the wealth of knowledge present at JSC. I wish I had come into the internship knowing exactly how helpful everyone in my branch would be, as I would have picked up on this sooner. I hope that having gained such a strong foundation in Simulink over this summer will open the opportunity to return to work on this project, or potentially other opportunities within the division. The idea of leaving a project I devoted ten weeks to is a hard one to cope with, so having the chance to pick up where I left off sounds appealing; alternatively, I am interested to see if there are any opening in the future that would allow me to work on a project that is more in-line with my research in estimation algorithms. Regardless, this summer has been a milestone in my professional career, and I hope this has started a long-term relationship between JSC and myself. I really enjoy the thought of building on my experience here over future summers while I work to complete my PhD at Missouri University of Science and Technology.

Gualdoni, Matthew J.↗

In-situ Nanoscale Ablation

An understanding of the ablation of carbon-based materials is crucial to modeling the behavior of atmospheric entry spacecrafts equipped with thermal protection systems (TPS).Carbon is the backbone of TPS systems such as PICA (phenolic-impregnated carbon ablator).Therefore in the present work we study ablation of highly oriented pyrolytic graphite (HOPG)in oxygen at temperatures up to 1000°C done within a gas reaction cell housed in a scanning transmission electron microscope (STEM). Observation of the HOPG oxidation on specimens sectioned parallel and normal to the carbon basal planes in the presence of oxygen are reported.Pitting caused by residual oxygen and/or platinum particles from the sectioning process was observed before oxygen gas flow was established at 750°C in the specimen sectioned with the carbon basal planes parallel to the beam direction. Introduction of oxygen flow caused rapid oxidation moving in a uniform front that completely consumed the HOPG in approximately 1.5minutes. The specimen sectioned with the carbon basal planes normal to the beam direction did not show the same pitting phenomenon but exhibited rapid oxidation at 1000°C that proceeded in a uniform front and completed in approximately 1 minute. All specimens tested had a husk resembling the original specimen shape left over after oxidation. It is concluded that this husk is most likely ash impurity from the HOPG or impurities from the sectioning process or E-chip. A residual gas analyszer (RGA) was successfully used to monitor gas flows during the experiments but was yet unsuccessful in monitoring oxidation gas products produced during the experiment.Future studies will optimize the RGA setup in order to maximize the potential for detecting these species. These in-situ studies successfully show how this highly ordered carbon ablates on a nano- and micro-scopic length scale and can be used to provide fundamental understanding of carbon ablation that can be used to design the next-generation of TPS systems.

Nanoscale Ablation↗

A Microfabricated Involute-Foil Regenerator for Stirling Engines

A segmented involute-foil regenerator has been designed, microfabricated and tested in an oscillating-flow rig with excellent results. During the Phase I effort, several approximations of parallel-plate regenerator geometry were chosen as potential candidates for a new microfabrication concept. Potential manufacturers and processes were surveyed. The selected concept consisted of stacked segmented-involute-foil disks (or annular portions of disks), originally to be microfabricated from stainless-steel via the LiGA (lithography, electroplating, and molding) process and EDM (electric discharge machining). During Phase II, re-planning of the effort led to test plans based on nickel disks, microfabricated via the LiGA process, only. A stack of nickel segmented-involute-foil disks was tested in an oscillating-flow test rig. These test results yielded a performance figure of merit (roughly the ratio of heat transfer to pressure drop) of about twice that of the 90% random fiber currently used in small ~ 100 W Stirling space-power convertors in the Reynolds Number range of interest (50-100). A Phase III effort is now underway to fabricate and test a segmented-involute-foil regenerator in a Stirling convertor. Though funding limitations prevent optimization of the Stirling engine geometry for use with this regenerator, the Sage computer code will be used to help evaluate the engine test results. Previous Sage Stirling model projections have indicated that a segmented-involute-foil regenerator is capable of improving the performance of an optimized involute-foil engine by 6-9%; it is also anticipated that such involute-foil geometries will be more reliable and easier to manufacture with tight-tolerance characteristics, than random-fiber or wire-screen regenerators. Beyond the near-term Phase III regenerator fabrication and engine testing, other goals are (1) fabrication from a material suitable for high temperature Stirling operation (up to 850 C for current engines; up to 1200 C for a potential engine-cooler for a Venus mission), and (2) reduction of the cost of the fabrication process to make it more suitable for terrestrial applications of segmented involute foils. Past attempts have been made to use wrapped foils to approximate the large theoretical figures of merit projected for parallel plates. Such metal wrapped foils have never proved very successful, apparently due to the difficulties of fabricating wrapped-foils with uniform gaps and maintaining the gaps under the stress of time-varying temperature gradients during start-up and shut-down, and relatively-steady temperature gradients during normal operation. In contrast, stacks of involute-foil disks, with each disk consisting of multiple involute-foil segments held between concentric circular ribs, have relatively robust structures. The oscillating-flow rig tests of the segmented-involute-foil regenerator have demonstrated a shift in regenerator performance strongly in the direction of the theoretical performance of ideal parallel-plate regenerators.

Tew, Roy↗

A Microfabricated Involute-Foil Regenerator for Stirling Engines

A segmented involute-foil regenerator has been designed, microfabricated and tested in an oscillating-flow rig with excellent results. During the Phase I effort, several approximations of parallel-plate regenerator geometry were chosen as potential candidates for a new microfabrication concept. Potential manufacturers and processes were surveyed. The selected concept consisted of stacked segmented-involute-foil disks (or annular portions of disks), originally to be microfabricated from stainless-steel via the LiGA (lithography, electroplating, and molding) process and EDM. During Phase II, re-planning of the effort led to test plans based on nickel disks, microfabricated via the LiGA process, only. A stack of nickel segmented-involute-foil disks was tested in an oscillating-flow test rig. These test results yielded a performance figure of merit (roughly the ratio of heat transfer to pressure drop) of about twice that of the 90 percent random fiber currently used in small approx.100 W Stirling space-power convertors-in the Reynolds Number range of interest (50 to 100). A Phase III effort is now underway to fabricate and test a segmented-involute-foil regenerator in a Stirling convertor. Though funding limitations prevent optimization of the Stirling engine geometry for use with this regenerator, the Sage computer code will be used to help evaluate the engine test results. Previous Sage Stirling model projections have indicated that a segmented-involute-foil regenerator is capable of improving the performance of an optimized involute-foil engine by 6 to 9 percent; it is also anticipated that such involute-foil geometries will be more reliable and easier to manufacture with tight-tolerance characteristics, than random-fiber or wire-screen regenerators. Beyond the near-term Phase III regenerator fabrication and engine testing, other goals are (1) fabrication from a material suitable for high temperature Stirling operation (up to 850 C for current engines; up to 1200 C for a potential engine-cooler for a Venus mission), and (2) reduction of the cost of the fabrication process to make it more suitable for terrestrial applications of segmented involute foils. Past attempts have been made to use wrapped foils to approximate the large theoretical figures of merit projected for parallel plates. Such metal wrapped foils have never proved very successful, apparently due to the difficulties of fabricating wrapped-foils with uniform gaps and maintaining the gaps under the stress of time-varying temperature gradients during start-up and shut-down, and relatively-steady temperature gradients during normal operation. In contrast, stacks of involute-foil disks, with each disk consisting of multiple involute-foil segments held between concentric circular ribs, have relatively robust structures. The oscillating-flow rig tests of the segmented-involute-foil regenerator have demonstrated a shift in regenerator performance strongly in the direction of the theoretical performance of ideal parallel-plate regenerators.

Tew, Roy↗

Integrating Cache Performance Modeling and Tuning Support in Parallelization Tools

With the resurgence of distributed shared memory (DSM) systems based on cache-coherent Non Uniform Memory Access (ccNUMA) architectures and increasing disparity between memory and processors speeds, data locality overheads are becoming the greatest bottlenecks in the way of realizing potential high performance of these systems. While parallelization tools and compilers facilitate the users in porting their sequential applications to a DSM system, a lot of time and effort is needed to tune the memory performance of these applications to achieve reasonable speedup. In this paper, we show that integrating cache performance modeling and tuning support within a parallelization environment can alleviate this problem. The Cache Performance Modeling and Prediction Tool (CPMP), employs trace-driven simulation techniques without the overhead of generating and managing detailed address traces. CPMP predicts the cache performance impact of source code level "what-if" modifications in a program to assist a user in the tuning process. CPMP is built on top of a customized version of the Computer Aided Parallelization Tools (CAPTools) environment. Finally, we demonstrate how CPMP can be applied to tune a real Computational Fluid Dynamics (CFD) application.

Waheed, Abdul↗

Parallelization of the Physical-Space Statistical Analysis System (PSAS)

Atmospheric data assimilation is a method of combining observations with model forecasts to produce a more accurate description of the atmosphere than the observations or forecast alone can provide. Data assimilation plays an increasingly important role in the study of climate and atmospheric chemistry. The NASA Data Assimilation Office (DAO) has developed the Goddard Earth Observing System Data Assimilation System (GEOS DAS) to create assimilated datasets. The core computational components of the GEOS DAS include the GEOS General Circulation Model (GCM) and the Physical-space Statistical Analysis System (PSAS). The need for timely validation of scientific enhancements to the data assimilation system poses computational demands that are best met by distributed parallel software. PSAS is implemented in Fortran 90 using object-based design principles. The analysis portions of the code solve two equations. The first of these is the "innovation" equation, which is solved on the unstructured observation grid using a preconditioned conjugate gradient (CG) method. The "analysis" equation is a transformation from the observation grid back to a structured grid, and is solved by a direct matrix-vector multiplication. Use of a factored-operator formulation reduces the computational complexity of both the CG solver and the matrix-vector multiplication, rendering the matrix-vector multiplications as a successive product of operators on a vector. Sparsity is introduced to these operators by partitioning the observations using an icosahedral decomposition scheme. PSAS builds a large (approx. 128MB) run-time database of parameters used in the calculation of these operators. Implementing a message passing parallel computing paradigm into an existing yet developing computational system as complex as PSAS is nontrivial. One of the technical challenges is balancing the requirements for computational reproducibility with the need for high performance. The problem of computational reproducibility is well known in the parallel computing community. It is a requirement that the parallel code perform calculations in a fashion that will yield identical results on different configurations of processing elements on the same platform. In some cases this problem can be solved by sacrificing performance. Meeting this requirement and still achieving high performance is very difficult. Topics to be discussed include: current PSAS design and parallelization strategy; reproducibility issues; load balance vs. database memory demands, possible solutions to these problems.

Larson, J. W.↗

Supercomputing on massively parallel bit-serial architectures

Research on the Goodyear Massively Parallel Processor (MPP) suggests that high-level parallel languages are practical and can be designed with powerful new semantics that allow algorithms to be efficiently mapped to the real machines. For the MPP these semantics include parallel/associative array selection for both dense and sparse matrices, variable precision arithmetic to trade accuracy for speed, micro-pipelined train broadcast, and conditional branching at the processing element (PE) control unit level. The preliminary design of a FORTRAN-like parallel language for the MPP has been completed and is being used to write programs to perform sparse matrix array selection, min/max search, matrix multiplication, Gaussian elimination on single bit arrays and other generic algorithms. A description is given of the MPP design. Features of the system and its operation are illustrated in the form of charts and diagrams.

Iobst, Ken↗