Search NASA⌕ Search

SEARCH · Search NASA

Results for “Performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Performance Evaluation in Network-Based Parallel Computing

Network-based parallel computing is emerging as a cost-effective alternative for solving many problems which require use of supercomputers or massively parallel computers. The primary objective of this project has been to conduct experimental research on performance evaluation for clustered parallel computing. First, a testbed was established by augmenting our existing SUNSPARCs' network with PVM (Parallel Virtual Machine) which is a software system for linking clusters of machines. Second, a set of three basic applications were selected. The applications consist of a parallel search, a parallel sort, a parallel matrix multiplication. These application programs were implemented in C programming language under PVM. Third, we conducted performance evaluation under various configurations and problem sizes. Alternative parallel computing models and workload allocations for application programs were explored. The performance metric was limited to elapsed time or response time which in the context of parallel computing can be expressed in terms of speedup. The results reveal that the overhead of communication latency between processes in many cases is the restricting factor to performance. That is, coarse-grain parallelism which requires less frequent communication between processes will result in higher performance in network-based computing. Finally, we are in the final stages of installing an Asynchronous Transfer Mode (ATM) switch and four ATM interfaces (each 155 Mbps) which will allow us to extend our study to newer applications, performance metrics, and configurations.

Dezhgosha, Kamyar↗

Performance Improvement Through Indexing of Turbine Airfoils: Numerical Simulation - Part 2

An experimental/analytical study has been conducted to determine the performance improvements achievable by circumferentially indexing succeeding rows of turbine stator airfoils. A series of tests was conducted to experimentally investigate stator wake clocking effects on the performance of the space shuttle main engine (SSME) alternate turbopump development (ATD) fuel turbine test article (TTA). The results from this study indicate that significant increases in stage efficiency can be attained through application of this airfoil clocking concept. Details of the experiment and its results are documented in part 1 of this paper. In order to gain insight into the mechanisms of the performance improvement, extensive computational fluid dynamics (CFD) simulations were executed. The subject of the present paper is the initial results from the CFD investigation of the configurations and conditions detailed in part 1 of the paper. To characterize the aerodynamic environments in the experimental test series, two-dimensional (2D), time accurate, multistage, viscous analyses were performed at the TTA midspan. Computational analyses for five different circumferential positions of the first stage stator have been completed. Details of the computational procedure and the results are presented. The analytical results verify the experimentally demonstrated performance improvement and are compared with data whenever possible. Predictions of time-averaged turbine efficiencies as well as gas conditions throughout the flow field are presented. An initial understanding of the turbine performance improvement mechanism based on the results from this investigation is described.

Griffin, Lisa W.↗

An Occupational Performance Test Validation Program for Fire Fighters at the Kennedy Space Center

We evaluated performance of a modified Combat Task Test (CTT) and of standard fitness tests in 20 male subjects to assess the prediction of occupational performance standards for Kennedy Space Center fire fighters. The CTT consisted of stair-climbing, a chopping simulation, and a victim rescue simulation. Average CTT performance time was 3.61 +/- 0.25 min (SEM) and all CTT tasks required 93% to 97% maximal heart rate. By using scores from the standard fitness tests, a multiple linear regression model was fitted to each parameter: the stairclimb (r(exp 2) = .905, P less than .05), the chopping performance time (r(exp 2) = .582, P less than .05), the victim rescue time (r(exp 2) = .218, P = not significant), and the total performance time (r(exp 2) = .769, P less than .05). Treadmill time was the predominant variable, being the major predictor in two of four models. These results indicated that standardized fitness tests can predict performance on some CTT tasks and that test predictors were amenable to exercise training.

Schonfeld, Brian R.↗

Advanced Technology Composite Fuselage-Structural Performance

Boeing is studying the technologies associated with the application of composite materials to commercial transport fuselage structure under the NASA-sponsored contracts for Advanced Technology Composite Aircraft Structures (ATCAS) and Materials Development Omnibus Contract (MDOC). This report addresses the program activities related to structural performance of the selected concepts, including both the design development and subsequent detailed evaluation. Design criteria were developed to ensure compliance with regulatory requirements and typical company objectives. Accurate analysis methods were selected and/or developed where practical, and conservative approaches were used where significant approximations were necessary. Design sizing activities supported subsequent development by providing representative design configurations for structural evaluation and by identifying the critical performance issues. Significant program efforts were directed towards assessing structural performance predictive capability. The structural database collected to perform this assessment was intimately linked to the manufacturing scale-up activities to ensure inclusion of manufacturing-induced performance traits. Mechanical tests were conducted to support the development and critical evaluation of analysis methods addressing internal loads, stability, ultimate strength, attachment and splice strength, and damage tolerance. Unresolved aspects of these performance issues were identified as part of the assessments, providing direction for future development.

Walker, T. H.↗

Adaptive Performance Seeking Control Using Fuzzy Model Reference Learning Control and Positive Gradient Control

Performance Seeking Control attempts to find the operating condition that will generate optimal performance and control the plant at that operating condition. In this paper a nonlinear multivariable Adaptive Performance Seeking Control (APSC) methodology will be developed and it will be demonstrated on a nonlinear system. The APSC is comprised of the Positive Gradient Control (PGC) and the Fuzzy Model Reference Learning Control (FMRLC). The PGC computes the positive gradients of the desired performance function with respect to the control inputs in order to drive the plant set points to the operating point that will produce optimal performance. The PGC approach will be derived in this paper. The feedback control of the plant is performed by the FMRLC. For the FMRLC, the conventional fuzzy model reference learning control methodology is utilized, with guidelines generated here for the effective tuning of the FMRLC controller.

Kopasakis, George↗

Towards High-Assurance High-Performance Program Synthesis

Domain-specific automatic program synthesis tools, also called application generators, are playing an ever-increasing role in software development. However, high-performance application generators require difficult manual construction, and are very difficult to verify correct. This paper describes research and an implemented system that transforms program synthesis tools based on deductive synthesis into high-performance application generators. Deductive synthesis uses theorem-proving to construct solutions when given problem specifications. The verification condition for a deductive synthesis tool is essentially the soundness of the implemented inference rules. Theory Operationalization for Program Synthesis (TOPS) synergistically combines reformulation, automated mathematical classification, and compilation through partial deduction to decision procedures. It transforms general-purpose deductive synthesis, with exponential performance, into efficient special-purpose deductive synthesis, with near-linear performance. This paper describes our experience with and empirical results of PD(TH) theory-based partial deduction - in which partial deduction of a set of first-order formulae is performed within the context of a background theory. The implemented TOPS system currently performs a special variant of PD(TH) in which the compilation process results in the transformation of a set of first order formulae into the theory of an instantiated library decision procedure augmented by a compiled unit theory.

Lowry, Michael↗

Portability and Cross-Platform Performance of an MPI-Based Parallel Polygon Renderer

Visualizing the results of computations performed on large-scale parallel computers is a challenging problem, due to the size of the datasets involved. One approach is to perform the visualization and graphics operations in place, exploiting the available parallelism to obtain the necessary rendering performance. Over the past several years, we have been developing algorithms and software to support visualization applications on NASA's parallel supercomputers. Our results have been incorporated into a parallel polygon rendering system called PGL. PGL was initially developed on tightly-coupled distributed-memory message-passing systems, including Intel's iPSC/860 and Paragon, and IBM's SP2. Over the past year, we have ported it to a variety of additional platforms, including the HP Exemplar, SGI Origin2OOO, Cray T3E, and clusters of Sun workstations. In implementing PGL, we have had two primary goals: cross-platform portability and high performance. Portability is important because (1) our manpower resources are limited, making it difficult to develop and maintain multiple versions of the code, and (2) NASA's complement of parallel computing platforms is diverse and subject to frequent change. Performance is important in delivering adequate rendering rates for complex scenes and ensuring that parallel computing resources are used effectively. Unfortunately, these two goals are often at odds. In this paper we report on our experiences with portability and performance of the PGL polygon renderer across a range of parallel computing platforms.

Crockett, Thomas W.↗

Demonstrating the Performance Benefits of the Strutjet RBCC for Space Launch Architectures

The Rocket Based Combined Cycle (RBCC) engine synergistically combines the best elements of airbreathing and rocket propulsion to benefit a wide range of future reusable launch vehicles (RLV). Aerojet's Strutjet RBCC offers high Isp during mid-phase acceleration, and high thrust for boost and final ascent phases. The result is a relatively low gross weight vehicle that reduces thrust requirements compared with all-rocket solutions. Relative to combination propulsion systems, the integrated propulsive elements of the Strutjet reduce engine weight and complexity. This paper will summarize the results of tests demonstrating our latest hydrogen-fueled Strutjet RBCC engine performance, including inlet operability and performance over a range of conditions, sea level static and Mach 2.4 rocket thrust augmentation, ramjet and scramjet performance, combined scramjet/rocket performance at Mach 8, and ascent mode rocket performance. These tests have significantly advanced the technology readiness of the Strutjet engine and substantiate the performance benefits of RBCC engines for reusable launch vehicle applications. A companion paper provides a focus on Strutjet as a basis for advanced air and space architecture, covering hydrogen and hydrocarbon fuels and cooled structures.

Siebenhaar, Adam↗

Rocket-in-a-Duct Performance Analysis

An axisymmetric, 110 N class, rocket configured with a free expansion between the rocket nozzle and a surrounding duct was tested in an altitude simulation facility. The propellants were gaseous hydrogen and gaseous oxygen and the hardware consisted of a heat sink type copper rocket firing through copper ducts of various diameters and lengths. A secondary flow of nitrogen was introduced at the blind end of the duct to mix with the primary rocket mass flow in the duct. This flow was in the range of 0 to 10% of the primary massflow and its effect on nozzle performance was measured. The random measurement errors on thrust and massflow were within +/-1%. One dimensional equilibrium calculations were used to establish the possible theoretical performance of these rocket-in-a-duct nozzles. Although the scale of these tests was small, they simulated the relevant flow expansion physics at a modest experimental cost. Test results indicated that lower performance was obtained at higher free expansion area ratios and longer ducts, while, higher performance was obtained with the addition of secondary flow. There was a discernable peak in specific impulse efficiency at 4% secondary flow. The small scale of these tests resulted in low performance efficiencies, but prior numerical modeling of larger rocket-in-a-duct engines predicted performance that was comparable to that of optimized rocket nozzles. This remains to be proven in large-scale, rocket-in-a-duct tests.

Schneider, Steven J.↗

Use of Boundary Layer Transition Detection to Validate Full-Scale Flight Performance Predictions

Full-scale flight performance predictions can be made using CFD or a combination of CFD and analytical skin-friction predictions. However, no matter what method is used to obtain full-scale flight performance predictions knowledge of the boundary layer state is critical. The implementation of CFD codes solving the Navier-Stokes equations to obtain these predictions is still a time consuming, expensive process. In addition, to ultimately obtain accurate performance predictions the transition location must be fixed in the CFD model. An example, using the M2.4-7A geometry, of the change in Navier-Stokes solution with changes in transition and in turbulence model will be shown. Oil flow visualization using the M2.4-7A 4.0% scale model in the 14'x22' wind tunnel shows that fixing transition at 10% x/c in the CFD model best captures the flow physics of the wing flow field. A less costly method of obtaining full-scale performance predictions is the use of non-linear Euler codes or linear CFD codes, such as panel methods, combined with analytical skin-friction predictions. Again, knowledge of the boundary layer state is critical to the accurate determination of full-scale flight performance. Boundary layer transition detection has been performed at 0.3 and 0.9 Mach numbers over an extensive Reynolds number range using the 2.2% scale Reference H model in the NTF. A temperature sensitive paint system was used to determine the boundary layer state for these conditions. Data was obtained for three configurations: the baseline, undeflected flaps configuration; the transonic cruise configuration; and, the high-lift configuration. It was determined that at low Reynolds number conditions, in the 8 to 10 million Reynolds number range, the baseline configuration has extensive regions of laminar flow, in fact significantly more than analytical skin-friction methods predict. This configuration is fully turbulent at about 30 million Reynolds number for both 0.3 and 0.9, Mach numbers. Both the transonic cruise and the high-lift configurations were fully turbulent aft of the leading-edge flap hingeline at all Reynolds numbers.

Hamner, Marvine↗

Performance Modeling and Measurement of Parallelized Code for Distributed Shared Memory Multiprocessors

This paper presents a model to evaluate the performance and overhead of parallelizing sequential code using compiler directives for multiprocessing on distributed shared memory (DSM) systems. With increasing popularity of shared address space architectures, it is essential to understand their performance impact on programs that benefit from shared memory multiprocessing. We present a simple model to characterize the performance of programs that are parallelized using compiler directives for shared memory multiprocessing. We parallelized the sequential implementation of NAS benchmarks using native Fortran77 compiler directives for an Origin2000, which is a DSM system based on a cache-coherent Non Uniform Memory Access (ccNUMA) architecture. We report measurement based performance of these parallelized benchmarks from four perspectives: efficacy of parallelization process; scalability; parallelization overhead; and comparison with hand-parallelized and -optimized version of the same benchmarks. Our results indicate that sequential programs can conveniently be parallelized for DSM systems using compiler directives but realizing performance gains as predicted by the performance model depends primarily on minimizing architecture-specific data locality overhead.

Waheed, Abdul↗

Predictors of Behavior and Performance in Extreme Environments: The Antarctic Space Analogue Program

To determine which, if any, characteristics should be incorporated into a select-in approach to screening personnel for long-duration spaceflight, we examined the influence of crewmember social/ demographic characteristics, personality traits, interpersonal needs, and characteristics of station physical environments on performance measures in 657 American men who spent an austral winter in Antarctica between 1963 and 1974. During screening, subjects completed a Personal History Questionnaire which obtained information on social and demographic characteristics, the Deep Freeze Opinion Survey which assessed 5 different personality traits, and the Fundamental Interpersonal Relations Orientation-Behavior (FIRO-B) Scale which measured 6 dimensions of interpersonal needs. Station environment included measures of crew size and severity of physical environment. Performance was assessed on the basis of combined peer-supervisor evaluations of overall performance, peer nominations of fellow crewmembers who made ideal winter-over candidates, and self-reported depressive symptoms. Social/demographic characteristics, personality traits, interpersonal needs, and characteristics of station environments collectively accounted for 9-17% of the variance in performance measures. The following characteristics were significant independent predictors of more than one performance measure: military service, low levels of neuroticism, extraversion and conscientiousness, and a low desire for affection from others. These results represent an important first step in the development of select-in criteria for personnel on long-duration missions in space and other extreme environments. These criteria must take into consideration the characteristics of the environment and the limitations they place on meeting needs for interpersonal relations and task performance, as well as the characteristics of the individuals and groups who live and work in these environments.

Palinkas, Lawrence A.↗

Applications Performance Under MPL and MPI on NAS IBM SP2

On July 5, 1994, an IBM Scalable POWER parallel System (IBM SP2) with 64 nodes, was installed at the Numerical Aerodynamic Simulation (NAS) Facility Each node of NAS IBM SP2 is a "wide node" consisting of a RISC 6000/590 workstation module with a clock of 66.5 MHz which can perform four floating point operations per clock with a peak performance of 266 Mflop/s. By the end of 1994, 64 nodes of IBM SP2 will be upgraded to 160 nodes with a peak performance of 42.5 Gflop/s. An overview of the IBM SP2 hardware is presented. The basic understanding of architectural details of RS 6000/590 will help application scientists the porting, optimizing, and tuning of codes from other machines such as the CRAY C90 and the Paragon to the NAS SP2. Optimization techniques such as quad-word loading, effective utilization of two floating point units, and data cache optimization of RS 6000/590 is illustrated, with examples giving performance gains at each optimization step. The conversion of codes using Intel's message passing library NX to codes using native Message Passing Library (MPL) and the Message Passing Interface (NMI) library available on the IBM SP2 is illustrated. In particular, we will present the performance of Fast Fourier Transform (FFT) kernel from NAS Parallel Benchmarks (NPB) under MPL and MPI. We have also optimized some of Fortran BLAS 2 and BLAS 3 routines, e.g., the optimized Fortran DAXPY runs at 175 Mflop/s and optimized Fortran DGEMM runs at 230 Mflop/s per node. The performance of the NPB (Class B) on the IBM SP2 is compared with the CRAY C90, Intel Paragon, TMC CM-5E, and the CRAY T3D.

Saini, Subhash↗

Isolated Performance at Mach Numbers From 0.60 to 2.86 of Several Expendable Nozzle Concepts for Supersonic Applications

Investigations have been conducted in the Langley 16-Foot Transonic Tunnel (at Mach numbers from 0.60 to 1.25) and in the Langley Unitary Plan Wind Tunnel (at Mach numbers from 2.16 to 2.86) at an angle of attack of 0 deg to determine the isolated performance of several expendable nozzle concepts for supersonic nonaugmented turbojet applications. The effects of centerbody base shape, shroud length, shroud ventilation, cruciform shroud expansion ratio, and cruciform shroud flap vectoring were investigated. The nozzle pressure ratio range, which was a function of Mach number, was between 1.9 and 11.8 in the 16-Foot Transonic Tunnel and between 7.9 and 54.9 in the Unitary Plan Wind Tunnel. Discharge coefficient, thrust-minus-drag, and the forces and moments generated by vectoring the divergent shroud flaps (for Mach numbers of 0.60 to 1.25 only) of a cruciform nozzle configuration were measured. The shortest nozzle had the best thrust-minus-drag performance at Mach numbers up to 0.95 but was approached in performance by other configurations at Mach numbers of 1.15 and 1.25. At Mach numbers above 1.25, the cruciform nozzle configuration having the same expansion ratio (2.64) as the fixed geometry nozzles had the best thrust-minus-drag performance. Ventilation of the fixed geometry divergent shrouds to the nozzle external boattail flow generally improved thrust-minus-drag performance at Mach numbers from 0.60 to 1.25, but decreased performance above a Mach number of 1.25.

Re, Richard J.↗

Performance Evaluation Methodologies and Tools for Massively Parallel Programs

The need for computing power has forced a migration from serial computation on a single processor to parallel processing on multiprocessors. However, without effective means to monitor (and analyze) program execution, tuning the performance of parallel programs becomes exponentially difficult as program complexity and machine size increase. The recent introduction of performance tuning tools from various supercomputer vendors (Intel's ParAide, TMC's PRISM, CSI'S Apprentice, and Convex's CXtrace) seems to indicate the maturity of performance tool technologies and vendors'/customers' recognition of their importance. However, a few important questions remain: What kind of performance bottlenecks can these tools detect (or correct)? How time consuming is the performance tuning process? What are some important technical issues that remain to be tackled in this area? This workshop reviews the fundamental concepts involved in analyzing and improving the performance of parallel and heterogeneous message-passing programs. Several alternative strategies will be contrasted, and for each we will describe how currently available tuning tools (e.g., AIMS, ParAide, PRISM, Apprentice, CXtrace, ATExpert, Pablo, IPS-2)) can be used to facilitate the process. We will characterize the effectiveness of the tools and methodologies based on actual user experiences at NASA Ames Research Center. Finally, we will discuss their limitations and outline recent approaches taken by vendors and the research community to address them.

Yan, Jerry C.↗

Performance Evaluation of Titanium Ion Optics for the NASA 30 cm Ion Thruster

The results of performance tests with titanium ion optics were presented and compared to those of molybdenum ion optics. Both titanium and molybdenum ion optics were initially operated until ion optics performance parameters achieved steady state values. Afterwards, performance characterizations were conducted. This permitted proper performance comparisons of titanium and molybdenum ion optics. Ion optics' performance A,as characterized over a broad thruster input power range of 0.5 to 3.0 kW. All performance parameters for titanium ion optics of achieved steady state values after processing 1200 gm of propellant. Molybdenum ion optics exhibited no burn-in. Impingement-limited total voltages for titanium ion optics where up to 55 V greater than those for molybdenum ion optics. Comparisons of electron backstreaming limits as a function of peak beam current density for molybdenum and titanium ion optics demonstrated that titanium ion optics operated with a higher electron backstreaming limit than molybdenum ion optics for a given peak beam current density. Screen grid ion transparencies for titanium ion optics were as much as 3.8 percent lower than those for molybdenum ion optics. Beam divergence half-angles that enclosed 95 percent of the total beam current for titanium ion optics were within 1 to 3 deg. of those for molybdenum ion optics. All beam divergence thrust correction factors for titanium ion optics were within 1 percent of those with molybdenum ion optics.

Soulas, George C.↗

Compute Server Performance Results

Parallel-vector supercomputers have been the workhorses of high performance computing. As expectations of future computing needs have risen faster than projected vector supercomputer performance, much work has been done investigating the feasibility of using Massively Parallel Processor systems as supercomputers. An even more recent development is the availability of high performance workstations which have the potential, when clustered together, to replace parallel-vector systems. We present a systematic comparison of floating point performance and price-performance for various compute server systems. A suite of highly vectorized programs was run on systems including traditional vector systems such as the Cray C90, and RISC workstations such as the IBM RS/6000 590 and the SGI R8000. The C90 system delivers 460 million floating point operations per second (FLOPS), the highest single processor rate of any vendor. However, if the price-performance ration (PPR) is considered to be most important, then the IBM and SGI processors are superior to the C90 processors. Even without code tuning, the IBM and SGI PPR's of 260 and 220 FLOPS per dollar exceed the C90 PPR of 160 FLOPS per dollar when running our highly vectorized suite,

Stockdale, I. E.↗

Performance Comparison of HPF and MPI Based NAS Parallel Benchmarks

Compilers supporting High Performance Form (HPF) features first appeared in late 1994 and early 1995 from Applied Parallel Research (APR), Digital Equipment Corporation, and The Portland Group (PGI). IBM introduced an HPF compiler for the IBM RS/6000 SP2 in April of 1996. Over the past two years, these implementations have shown steady improvement in terms of both features and performance. The performance of various hardware/ programming model (HPF and MPI) combinations will be compared, based on latest NAS Parallel Benchmark results, thus providing a cross-machine and cross-model comparison. Specifically, HPF based NPB results will be compared with MPI based NPB results to provide perspective on performance currently obtainable using HPF versus MPI or versus hand-tuned implementations such as those supplied by the hardware vendors. In addition, we would also present NPB, (Version 1.0) performance results for the following systems: DEC Alpha Server 8400 5/440, Fujitsu CAPP Series (VX, VPP300, and VPP700), HP/Convex Exemplar SPP2000, IBM RS/6000 SP P2SC node (120 MHz), NEC SX-4/32, SGI/CRAY T3E, and SGI Origin2000. We would also present sustained performance per dollar for Class B LU, SP and BT benchmarks.

Saini, Subhash↗