Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Performance Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Use Computer-Aided Tools to Parallelize Large CFD Applications

Porting applications to high performance parallel computers is always a challenging task. It is time consuming and costly. With rapid progressing in hardware architectures and increasing complexity of real applications in recent years, the problem becomes even more sever. Today, scalability and high performance are mostly involving handwritten parallel programs using message-passing libraries (e.g. MPI). However, this process is very difficult and often error-prone. The recent reemergence of shared memory parallel (SMP) architectures, such as the cache coherent Non-Uniform Memory Access (ccNUMA) architecture used in the SGI Origin 2000, show good prospects for scaling beyond hundreds of processors. Programming on an SMP is simplified by working in a globally accessible address space. The user can supply compiler directives, such as OpenMP, to parallelize the code. As an industry standard for portable implementation of parallel programs for SMPs, OpenMP is a set of compiler directives and callable runtime library routines that extend Fortran, C and C++ to express shared memory parallelism. It promises an incremental path for parallel conversion of existing software, as well as scalability and performance for a complete rewrite or an entirely new development. Perhaps the main disadvantage of programming with directives is that inserted directives may not necessarily enhance performance. In the worst cases, it can create erroneous results. While vendors have provided tools to perform error-checking and profiling, automation in directive insertion is very limited and often failed on large programs, primarily due to the lack of a thorough enough data dependence analysis. To overcome the deficiency, we have developed a toolkit, CAPO, to automatically insert OpenMP directives in Fortran programs and apply certain degrees of optimization. CAPO is aimed at taking advantage of detailed inter-procedural dependence analysis provided by CAPTools, developed by the University of Greenwich, to reduce potential errors made by users. Earlier tests on NAS Benchmarks and ARC3D have demonstrated good success of this tool. In this study, we have applied CAPO to parallelize three large applications in the area of computational fluid dynamics (CFD): OVERFLOW, TLNS3D and INS3D. These codes are widely used for solving Navier-Stokes equations with complicated boundary conditions and turbulence model in multiple zones. Each one comprises of from 50K to 1,00k lines of FORTRAN77. As an example, CAPO took 77 hours to complete the data dependence analysis of OVERFLOW on a workstation (SGI, 175MHz, R10K processor). A fair amount of effort was spent on correcting false dependencies due to lack of necessary knowledge during the analysis. Even so, CAPO provides an easy way for user to interact with the parallelization process. The OpenMP version was generated within a day after the analysis was completed. Due to sequential algorithms involved, code sections in TLNS3D and INS3D need to be restructured by hand to produce more efficient parallel codes. An included figure shows preliminary test results of the generated OVERFLOW with several test cases in single zone. The MPI data points for the small test case were taken from a handcoded MPI version. As we can see, CAPO's version has achieved 18 fold speed up on 32 nodes of the SGI O2K. For the small test case, it outperformed the MPI version. These results are very encouraging, but further work is needed. For example, although CAPO attempts to place directives on the outer- most parallel loops in an interprocedural framework, it does not insert directives based on the best manual strategy. In particular, it lacks the support of parallelization at the multi-zone level. Future work will emphasize on the development of methodology to work in a multi-zone level and with a hybrid approach. Development of tools to perform more complicated code transformation is also needed.

Jin, H.↗

An interactive parallel programming environment applied in atmospheric science

This article introduces an interactive parallel programming environment (IPPE) that simplifies the generation and execution of parallel programs. One of the tasks of the environment is to generate message-passing parallel programs for homogeneous and heterogeneous computing platforms. The parallel programs are represented by using visual objects. This is accomplished with the help of a graphical programming editor that is implemented in Java and enables portability to a wide variety of computer platforms. In contrast to other graphical programming systems, reusable parts of the programs can be stored in a program library to support rapid prototyping. In addition, runtime performance data on different computing platforms is collected in a database. A selection process determines dynamically the software and the hardware platform to be used to solve the problem in minimal wall-clock time. The environment is currently being tested on a Grand Challenge problem, the NASA four-dimensional data assimilation system.

Interactive Display Devices↗

Load balancing and task decomposition techniques for parallel implementation of integrated vision systems algorithms

Several techniques are presented to perform static and dynamic load balancing schemes for integrated vision systems. These techniques are novel in the sense that they capture the computational requirements of a task by examining the data when they are produced. Furthermore, they can be applied to many integrated vision systems because many algorithms in different systems are either the same or have similar computational characteristics. These techniques are evaluated by applying them to the algorithms in a motion estimation system. It is shown that the performance gains when these techniques are used are significant and the overhead of using these techniques is minimal. The performance is evaluated by implementing the algorithms using the presented techniques on a hypercube multiprocessor system.

Choudhary, Alok N.↗

An examination of injector/combustor design effects on scramjet performance

Description of a simplified one-dimensional treatment of fuel injection for supersonic combustor performance analysis. Representative mixing efficiency variations for both parallel and cross-stream injection are obtained, and approximate means for estimating the length required for complete mixing are demonstrated. Comparisons of calculated and measured data show good agreement.

Anderson, G. Y.↗

Experimental investigation of a swept-strut fuel-injector concept for scramjet application

Results are presented of an experiment to investigate the behavior at Mach 4 flight conditions of the swept-strut fuel-injector concept employed in the Langley integrated modular scramjet engine design. Autoignition of the hydrogen fuel was not achieved at stagnation temperatures corresponding to a flight Mach number of 4; however, once ignition was achieved, stable combustion was maintained. Pressure disturbances upstream of the injector location, which were caused by fuel injection and combustion, were generally not observed; this indicates the absence of serious adverse combustor-inlet interactions. Mixing performance and reaction performance determined from probe surveys and wall pressure data indicate that high combustion efficiency should be obtained with the combustor length provided in the scramjet engine design. No adverse interaction between the perpendicular and parallel fuel-injection modes was observed.

Anderson, G. Y.↗

High temperature, low mass solar blanket development

This paper presents methods of incorporating ultrathin silicon solar cells into photovoltaic blankets for space applications. This type of cell has the highest power-to-mass ratio and best performance under space radiation of any silicon solar cell. Interconnect materials and designs, and the results of the investigation of the applicability of parallel-gap resistance welding for interconnecting ultrathin cells are discussed. Data relating contact pull strength and cell electrical degradation to welding parameters such as time, voltage, and pressure are presented. Methods for bonding ultrathin cells to flexible substrates and for bonding thin covers to these cells are described, and the results of vacuum thermal cycling and thermal soak tests on prototype ultrathin cell test coupons are included.

Mesch, H. G.↗

Crossed orbit interferometry - Theory and experimental results from SIR-B

Crossed orbit interferometry, which can perform measurements with only one antenna making two images of a scene during two separate passes and can operate even if the orbits are not parallel, is discussed and tested using SIR-B data. It is found that a Doppler refocusing of the SAR azimuth correlation, involving a resampling of one of the imgages in the cross-track direction, is necessary to remove the linear shift of the scene. The refocusing process also involves a terrain dependent resampling in the azimuth direction. A method for finding tie points to guide the resampling is discussed and a coarse altitude map derived only from the tie points is presented. Spatial heterodyned interferograms that contain the effects of the crossed orbit geometry are presented and a theoretical model is developed to explain them. The model is extended to calculate altitudes from the interferograms, and a final altitude map is presented.

Gabriel, Andrew K.↗

Highly-Parallel, Highly-Compact Computing Structures Implemented in Nanotechnology

In this paper, we describe work in which we are evaluating how the evolving properties of nano-electronic devices could best be utilized in highly parallel computing structures. Because of their combination of high performance, low power, and extreme compactness, such structures would have obvious applications in spaceborne environments, both for general mission control and for on-board data analysis. However, the anticipated properties of nano-devices mean that the optimum architecture for such systems is by no means certain. Candidates include single instruction multiple datastream (SIMD) arrays, neural networks, and multiple instruction multiple datastream (MIMD) assemblies.

Crawley, D. G.↗

Performance of the Galley Parallel File System

As the input/output (I/O) needs of parallel scientific applications increase, file systems for multiprocessors are being designed to provide applications with parallel access to multiple disks. Many parallel file systems present applications with a conventional Unix-like interface that allows the application to access multiple disks transparently. This interface conceals the parallism within the file system, which increases the ease of programmability, but makes it difficult or impossible for sophisticated programmers and libraries to use knowledge about their I/O needs to exploit that parallelism. Furthermore, most current parallel file systems are optimized for a different workload than they are being asked to support. We introduce Galley, a new parallel file system that is intended to efficiently support realistic parallel workloads. Initial experiments, reported in this paper, indicate that Galley is capable of providing high-performance 1/O to applications the applications that rely on them. In Section 3 we describe that access data in patterns that have been observed to be common.

Nieuwejaar, Nils↗

"Genetically Engineered" Nanoelectronics

The quantum mechanical functionality of nanoelectronic devices such as resonant tunneling diodes (RTDs), quantum well infrared-photodetectors (QWIPs), quantum well lasers, and heterostructure field effect transistors (HFETs) is enabled by material variations on an atomic scale. The design and optimization of such devices requires a fundamental understanding of electron transport in such dimensions. The Nanoelectronic Modeling Tool (NEMO) is a general-purpose quantum device design and analysis tool based on a fundamental non-equilibrium electron transport theory. NEW was combined with a parallelized genetic algorithm package (PGAPACK) to evolve structural and material parameters to match a desired set of experimental data. A numerical experiment that evolves structural variations such as layer widths and doping concentrations is performed to analyze an experimental current voltage characteristic. The genetic algorithm is found to drive the NEMO simulation parameters close to the experimentally prescribed layer thicknesses and doping profiles. With such a quantitative agreement between theory and experiment design synthesis can be performed.

Klimeck, Gerhard↗

Proteomic Retrieval from Nucleic Acid Depleted Space-Flown Human Cells

Compared to experiments utilizing humans in microgravity, cell-based approaches to questions about subsystems of the human system afford multiple advantages, such as crew safety and the ability to achieve statistical significance. To maximize the science return from flight samples, an optimized method was developed to recover protein from samples depleted of nucleic acid. This technique allows multiple analyses on a single cellular sample and when applied to future cellular investigations could accelerate solutions to significant biomedical barriers to human space exploration. Cell cultures grown in American Fluoroseal bags were treated with an RNA stabilizing agent (RNAlater - Ambion), which enabled both RNA and immunoreactive protein analyses. RNA was purified using an RNAqueous(registered TradeMark) kit (Ambion) and the remaining RNA free supernatant was precipitated with 5% trichloroacetic acid. The precipitate was dissolved in SDS running buffer and tested for protein content using a bicinchoninic acid assay (1) (Sigma). Equal loads of protein were placed on SDS-PAGE gels and either stained with CyproOrange (Amersham) or transferred using Western Blotting techniques (2,3,4). Protein recovered from RNAlater-treated cells and stained with protein stain, was measured using Imagequant volume measurements for rectangles of equal size. BSA treated in this way gave quantitative data over the protein range used (Fig 1). Human renal cortical epithelial (HRCE) cells (5,6,7) grown onboard the International Space Station (ISS) during Increment 3 and in ground control cultures exhibited similar immunoreactivity profiles for antibodies to the Vitamin D receptor (VDR) (Fig 2), the beta isoform of protein kinase C (PKC ) (Fig 3), and glyceraldehyde-3-phosphate dehydrogenase (GAPDH) (Fig 4). Parallel immunohistochemical studies on formalin-fixed flight and ground control cultures also showed positive immunostaining for VDR and other biomarkers (Fig 5). These results are consistent with data from additional antigenic recovery experiments performed on human Mullerian tumor cells cultured in microgravity (8).

Hammond, D. K.↗

The Design and Evaluation of "CAPTools"--A Computer Aided Parallelization Toolkit

Writing applications for high performance computers is a challenging task. Although writing code by hand still offers the best performance, it is extremely costly and often not very portable. The Computer Aided Parallelization Tools (CAPTools) are a toolkit designed to help automate the mapping of sequential FORTRAN scientific applications onto multiprocessors. CAPTools consists of the following major components: an inter-procedural dependence analysis module that incorporates user knowledge; a 'self-propagating' data partitioning module driven via user guidance; an execution control mask generation and optimization module for the user to fine tune parallel processing of individual partitions; a program transformation/restructuring facility for source code clean up and optimization; a set of browsers through which the user interacts with CAPTools at each stage of the parallelization process; and a code generator supporting multiple programming paradigms on various multiprocessors. Besides describing the rationale behind the architecture of CAPTools, the parallelization process is illustrated via case studies involving structured and unstructured meshes. The programming process and the performance of the generated parallel programs are compared against other programming alternatives based on the NAS Parallel Benchmarks, ARC3D and other scientific applications. Based on these results, a discussion on the feasibility of constructing architectural independent parallel applications is presented.

Yan, Jerry↗

Operations automation

This is truly the era of 'faster-better-cheaper' at the National Aeronautics and Space Administration/Jet Propulsion Laboratory (NASA/JPL). To continue JPL's primary mission of building and operating interplanetary spacecraft, all possible avenues are being explored in the search for better value for each dollar spent. A significant cost factor in any mission is the amount of manpower required to receive, decode, decommutate, and distribute spacecraft engineering and experiment data. The replacement of the many mission-unique data systems with the single Advanced Multimission Operations System (AMMOS) has already allowed for some manpower reduction. Now, we find that further economies are made possible by drastically reducing the number of human interventions required to perform the setup, data saving, station handover, processed data loading, and tear down activities that are associated with each spacecraft tracking pass. We have recently adapted three public domain tools to the AMMOS system which allow common elements to be scheduled and initialized without the normal human intervention. This is accomplished with a stored weekly event schedule. The manual entries and specialized scripts which had to be provided just prior to and during a pass are now triggered by the schedule to perform the functions unique to the upcoming pass. This combination of public domain software and the AMMOS system has been run in parallel with the flight operation in an online testing phase for six months. With this methodology, a savings of 11 man-years per year is projected with no increase in data loss or project risk. There are even greater savings to be gained as we learn other uses for this configuration.

Boreham, Charles Thomas↗

Post-analysis report on Chesapeake Bay data processing

The additional processing performed on data collected over the Rhode River Test Site and Forestry Site in November 1970 is reported. The techniques and procedures used to obtain the processed results are described. Thermal data collected over three approximately parallel lines of the site were contoured, and the results color coded, for the purpose of delineating important scene constituents and to identify trees attacked by pine bark beetles. Contouring work and histogram preparation are reviewed and the important conclusions from the spectral analysis and recognition computer (SPARC) signature extension work are summarized. The SPARC setup and processing records are presented and recommendations are made for future data collection over the site.

Thomson, F.↗

Performance of battery reconditioning on the communications technology satellite

Some real life data on the flight performance and reconditioning on the Communications Technology Satellite are presented. There are two 24-cell nominal 5 amp-hour nickel cadmium batteries on board. The two batteries are operated in parallel with isolating diodes in between and permanently connected to the housekeeping buses. The recharge can be at C over 10 or C over 20 to a 1.4 charge/discharge ratio; and recharge is terminated either when a computed 1.4 charge/discharge ratio is reached, or a charge voltage peak is reached.

Lackner, J.↗

The NAS parallel benchmarks

A new set of benchmarks has been developed for the performance evaluation of highly parallel supercomputers in the framework of the NASA Ames Numerical Aerodynamic Simulation (NAS) Program. These consist of five 'parallel kernel' benchmarks and three 'simulated application' benchmarks. Together they mimic the computation and data movement characteristics of large-scale computational fluid dynamics applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification-all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, D. H.↗

An analysis of I/O efficient order-statistic-based techniques for noise power estimation in the HRMS sky survey's operational system

Noise power estimation in the High-Resolution Microwave Survey (HRMS) sky survey element is considered as an example of a constant false alarm rate (CFAR) signal detection problem. Order-statistic-based noise power estimators for CFAR detection are considered in terms of required estimator accuracy and estimator dynamic range. By limiting the dynamic range of the value to be estimated, the performance of an order-statistic estimator can be achieved by simpler techniques requiring only a single pass of the data. Simple threshold-and-count techniques are examined, and it is shown how several parallel threshold-and-count estimation devices can be used to expand the dynamic range to meet HRMS system requirements with minimal hardware complexity. An input/output (I/O) efficient limited-precision order-statistic estimator with wide but limited dynamic range is also examined.

Zimmerman, G. A.↗

Lead-Free Experiment in a Space Environment

This Technical Memorandum addresses the Lead-Free Technology Experiment in Space Environment that flew as part of the seventh Materials International Space Station Experiment outside the International Space Station for approximately 18 months. Its intent was to provide data on the performance of lead-free electronics in an actual space environment. Its postflight condition is compared to the preflight condition as well as to the condition of an identical package operating in parallel in the laboratory. Some tin whisker growth was seen on a flight board but the whiskers were few and short. There were no solder joint failures, no tin pest formation, and no significant intermetallic compound formation or growth on either the flight or ground units.

Blanche, J. F.↗