Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel hybrid”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Parallel Unsteady Turbopump Simulations for Liquid Rocket Engines

This paper reports the progress being made towards complete turbo-pump simulation capability for liquid rocket engines. Space Shuttle Main Engine (SSME) turbo-pump impeller is used as a test case for the performance evaluation of the MPI and hybrid MPI/Open-MP versions of the INS3D code. Then, a computational model of a turbo-pump has been developed for the shuttle upgrade program. Relative motion of the grid system for rotor-stator interaction was obtained by employing overset grid techniques. Time-accuracy of the scheme has been evaluated by using simple test cases. Unsteady computations for SSME turbo-pump, which contains 136 zones with 35 Million grid points, are currently underway on Origin 2000 systems at NASA Ames Research Center. Results from time-accurate simulations with moving boundary capability, and the performance of the parallel versions of the code will be presented in the final paper.

Kiris, Cetin C.↗

Electronic hardware implementations of neutral networks

This paper examines some of the present work on the development of electronic neural network hardware. In particular, the investigations currently under way at JPL on neural network hardware implementations based on custom VLSI technology, novel thin film materials, and an analog-digital hybrid architecture are reviewed. The availability of such hardware will greatly benefit and enhance the present intense research effort on the potential computational capabilities of highly parallel systems based on neural network models.

Thakoor, A. P.↗

Emissions Prediction and Measurement for Liquid-Fueled TVC Combustor with and without Water Injection

An investigation is performed to evaluate the performance of a computational fluid dynamics (CFD) tool for the prediction of the reacting flow in a liquid-fueled combustor that uses water injection for control of pollutant emissions. The experiment consists of a multisector, liquid-fueled combustor rig operated at different inlet pressures and temperatures, and over a range of fuel/air and water/fuel ratios. Fuel can be injected directly into the main combustion airstream and into the cavities. Test rig performance is characterized by combustor exit quantities such as temperature and emissions measurements using rakes and overall pressure drop from upstream plenum to combustor exit. Visualization of the flame is performed using gray scale and color still photographs and high-frame-rate videos. CFD simulations are performed utilizing a methodology that includes computer-aided design (CAD) solid modeling of the geometry, parallel processing over networked computers, and graphical and quantitative post-processing. Physical models include liquid fuel droplet dynamics and evaporation, with combustion modeled using a hybrid finite-rate chemistry model developed for Jet-A fuel. CFD and experimental results are compared for cases with cavity-only fueling, while numerical studies of cavity and main fueling was also performed. Predicted and measured trends in combustor exit temperature, CO and NOx are in general agreement at the different water/fuel loading rates, although quantitative differences exist between the predictions and measurements.

Brankovic, A.↗

Parallel Monotonic Basin Hopping for Low Thrust Trajectory Optimization

Monotonic Basin Hopping has been shown to be an effective method of solving low thrust trajectory optimization problems. This paper outlines an extension to the common serial implementation by parallelizing it over any number of available compute cores. The Parallel Monotonic Basin Hopping algorithm described herein is shown to be an effective way to more quickly locate feasible solutions, and improve locally optimal solutions in an automated way without requiring a feasible initial guess. The increased speed achieved through parallelization enables the algorithm to be applied to more complex problems that would otherwise be impractical for a serial implementation. Low thrust cislunar transfers and a hybrid Mars example case demonstrate the effectiveness of the algorithm. Finally, a preliminary scaling study quantifies the expected decrease in solve time compared to a serial implementation.,

McCarty, Steven L.↗

Parallel Monotonic Basin Hopping for Low Thrust Trajectory Optimization

Monotonic Basin Hopping has been shown to be an effective method of solving low thrust trajectory optimization problems. This paper outlines an extension to the common serial implementation by parallelizing it over any number of available compute cores. The Parallel Monotonic Basin Hopping algorithm described herein is shown to be an effective way to more quickly locate feasible solutions, and improve locally optimal solutions in an automated way without requiring a feasible initial guess. The increased speed achieved through parallelization enables the algorithm to be applied to more complex problems that would otherwise be impractical for a serial implementation. Low thrust cislunar transfers and a hybrid Mars example case demonstrate the effectiveness of the algorithm. Finally, a preliminary scaling study quantifies the expected decrease in solve time compared to a serial implementation.

McCarty, Steven L.↗

Architecture and design of a 500-MHz gallium-arsenide processing element for a parallel supercomputer

The design of the processing element of GASP, a GaAs supercomputer with a 500-MHz instruction issue rate and 1-GHz subsystem clocks, is presented. The novel, functionally modular, block data flow architecture of GASP is described. The architecture and design of a GASP processing element is then presented. The processing element (PE) is implemented in a hybrid semiconductor module with 152 custom GaAs ICs of eight different types. The effects of the implementation technology on both the system-level architecture and the PE design are discussed. SPICE simulations indicate that parts of the PE are capable of being clocked at 1 GHz, while the rest of the PE uses a 500-MHz clock. The architecture utilizes data flow techniques at a program block level, which allows efficient execution of parallel programs while maintaining reasonably good performance on sequential programs. A simulation study of the architecture indicates that an instruction execution rate of over 30,000 MIPS can be attained with 65 PEs.

Fouts, Douglas J.↗

Hybrid Particle-Element Simulation of Impact on Composite Orbital Debris Shields

This report describes the development of new numerical methods and new constitutive models for the simulation of hypervelocity impact effects on spacecraft. The research has included parallel implementation of the numerical methods and material models developed under the project. Validation work has included both one dimensional simulations, for comparison with exact solutions, and three dimensional simulations of published hypervelocity impact experiments. The validated formulations have been applied to simulate impact effects in a velocity and kinetic energy regime outside the capabilities of current experimental methods. The research results presented here allow for the expanded use of numerical simulation, as a complement to experimental work, in future design of spacecraft for hypervelocity impact effects.

Fahrenthold, Eric P.↗

Performance of a 300 Mbps 1:16 serial/parallel optoelectronic receiver module

Optical interconnects are being considered for the high speed distribution of multiplexed control signals in GaAs monolithic microwave integrated circuit (MMIC) based phased array antennas. The performance of a hybrid GaAs optoelectronic integrated circuit (OEIC) is described, as well as its design and fabrication. The OEIC converts a 16-bit serial optical input to a 16 parallel line electrical output using an on-board 1:16 demultiplexer and operates at data rates as high as 30b Mbps. The performance characteristics and potential applications of the device are presented.

Richard, M. A.↗

Performance of a 300 Mbps 1:16 serial/parallel optoelectronic receiver module

Optical interconnects are being considered for the high speed distribution of multiplexed control signals in GaAs MMIC-based phased array antennas. This paper describes the performance of a hybrid GaAs optoelectronic integrated circuit (OEIC), along with a description of its design and fabrication. The OEIC converts a 16-bit serial optical input to a 16 parallel line electrical output using an on-board 1:16 demultiplexer and operates at data rates as high as 305 Mbps. The performance characteristics as well as potential applications of the device are presented.

Richard, M. A.↗

Use Computer-Aided Tools to Parallelize Large CFD Applications

Porting applications to high performance parallel computers is always a challenging task. It is time consuming and costly. With rapid progressing in hardware architectures and increasing complexity of real applications in recent years, the problem becomes even more sever. Today, scalability and high performance are mostly involving handwritten parallel programs using message-passing libraries (e.g. MPI). However, this process is very difficult and often error-prone. The recent reemergence of shared memory parallel (SMP) architectures, such as the cache coherent Non-Uniform Memory Access (ccNUMA) architecture used in the SGI Origin 2000, show good prospects for scaling beyond hundreds of processors. Programming on an SMP is simplified by working in a globally accessible address space. The user can supply compiler directives, such as OpenMP, to parallelize the code. As an industry standard for portable implementation of parallel programs for SMPs, OpenMP is a set of compiler directives and callable runtime library routines that extend Fortran, C and C++ to express shared memory parallelism. It promises an incremental path for parallel conversion of existing software, as well as scalability and performance for a complete rewrite or an entirely new development. Perhaps the main disadvantage of programming with directives is that inserted directives may not necessarily enhance performance. In the worst cases, it can create erroneous results. While vendors have provided tools to perform error-checking and profiling, automation in directive insertion is very limited and often failed on large programs, primarily due to the lack of a thorough enough data dependence analysis. To overcome the deficiency, we have developed a toolkit, CAPO, to automatically insert OpenMP directives in Fortran programs and apply certain degrees of optimization. CAPO is aimed at taking advantage of detailed inter-procedural dependence analysis provided by CAPTools, developed by the University of Greenwich, to reduce potential errors made by users. Earlier tests on NAS Benchmarks and ARC3D have demonstrated good success of this tool. In this study, we have applied CAPO to parallelize three large applications in the area of computational fluid dynamics (CFD): OVERFLOW, TLNS3D and INS3D. These codes are widely used for solving Navier-Stokes equations with complicated boundary conditions and turbulence model in multiple zones. Each one comprises of from 50K to 1,00k lines of FORTRAN77. As an example, CAPO took 77 hours to complete the data dependence analysis of OVERFLOW on a workstation (SGI, 175MHz, R10K processor). A fair amount of effort was spent on correcting false dependencies due to lack of necessary knowledge during the analysis. Even so, CAPO provides an easy way for user to interact with the parallelization process. The OpenMP version was generated within a day after the analysis was completed. Due to sequential algorithms involved, code sections in TLNS3D and INS3D need to be restructured by hand to produce more efficient parallel codes. An included figure shows preliminary test results of the generated OVERFLOW with several test cases in single zone. The MPI data points for the small test case were taken from a handcoded MPI version. As we can see, CAPO's version has achieved 18 fold speed up on 32 nodes of the SGI O2K. For the small test case, it outperformed the MPI version. These results are very encouraging, but further work is needed. For example, although CAPO attempts to place directives on the outer- most parallel loops in an interprocedural framework, it does not insert directives based on the best manual strategy. In particular, it lacks the support of parallelization at the multi-zone level. Future work will emphasize on the development of methodology to work in a multi-zone level and with a hybrid approach. Development of tools to perform more complicated code transformation is also needed.

Jin, H.↗

A new environment to simulate the dynamics in the close proximity of rubble-pile asteroids

This paper presents a new environment to simulate close-proximity dynamics around rubble-pile asteroids. The code provides methods for modeling the asteroid’s gravity field and surface through granular dynamics. It implements stateof-the-art techniques to model both gravity and contact interaction between particles: 1) mutual gravity as either direct N2 or Barnes-Hut GPU-parallel octree and 2) contact dynamics with a soft-body (force-based, smooth dynamics), hard-body (constraint-based, non-smooth dynamics), or hybrid (constraint-based with compliance and damping) approach. A very relevant feature of the code is its ability to handle complex-shaped rigid bodies and their full 6D motion. Examples of spacecraft close-proximity scenarios and their numerical simulations are shown.

Ferrari, Fabio↗

Hybrid VLSI/QCA Architecture for Computing FFTs

A data-processor architecture that would incorporate elements of both conventional very-large-scale integrated (VLSI) circuitry and quantum-dot cellular automata (QCA) has been proposed to enable the highly parallel and systolic computation of fast Fourier transforms (FFTs). The proposed circuit would complement the QCA-based circuits described in several prior NASA Tech Briefs articles, namely Implementing Permutation Matrices by Use of Quantum Dots (NPO-20801), Vol. 25, No. 10 (October 2001), page 42; Compact Interconnection Networks Based on Quantum Dots (NPO-20855) Vol. 27, No. 1 (January 2003), page 32; and Bit-Serial Adder Based on Quantum Dots (NPO-20869), Vol. 27, No. 1 (January 2003), page 35. The cited prior articles described the limitations of very-large-scale integrated (VLSI) circuitry and the major potential advantage afforded by QCA. To recapitulate: In a VLSI circuit, signal paths that are required not to interact with each other must not cross in the same plane. In contrast, for reasons too complex to describe in the limited space available for this article, suitably designed and operated QCAbased signal paths that are required not to interact with each other can nevertheless be allowed to cross each other in the same plane without adverse effect. In principle, this characteristic could be exploited to design compact, coplanar, simple (relative to VLSI) QCA-based networks to implement complex, advanced interconnection schemes.

Fijany, Amir↗

Nonlinear generation of whistler waves by an ion beam

An electromagnetic hybrid code is used to simulate a new mechanism for whistler wave generation by an ion beam. First, a field-aligned ion beam becomes unstable to the electromagnetic ion/ion right-hand resonant instability which generates large amplitude MHD-like waves. These waves then trap the ion beam and increase its effective temperature anisotropy. As a result, the growth rates of the electron/whistler instability are significantly enhanced, and whistlers start to grow above the noise level. At the same time, because of the reduced parallel drift speed of the ion beam, the frequencies of the whistlers are also downshifted. Full simulations were performed to isolate and separately investigate the electron/ion whistler instability. The results are in agreement with the assumption of fluid electrons in the hybrid simulations and with the linear theory of the instability.

Akimoto, K.↗

Boeing Helicopters Advanced Rotorcraft Transmission (ART) Program summary of component tests

The principal objectives of the ART program are briefly reviewed, and the results of advanced technology component tests are summarized. The tests discussed include noise reduction by active cancellation, hybrid bidirectional tapered roller bearings, improved bearing life theory and friction tests, transmission lube study with hybrid bearings, and precision near-net-shape forged spur gears. Attention is also given to the study of high profile contact ratio noninvolute tooth form spur gears, parallel axis gear noise study, and surface modified titanium accessory spur gears.

Lenski, Joseph W., Jr.↗

Test Bed Doppler Wind Lidar and Intercomparison Facility At NASA Langley Research Center

State of the art 2-micron lasers and other lidar components under development by NASA are being demonstrated and validated in a mobile test bed Doppler wind lidar. A lidar intercomparison facility has been developed to ensure parallel alignment of up to 4 Doppler lidar systems while measuring wind. Investigations of the new components; their operation in a complete system; systematic and random errors; the hybrid (joint coherent and direct detection) approach to global wind measurement; and atmospheric wind behavior are planned. Future uses of the VALIDAR (VALIDation LIDAR) mobile lidar may include comparison with the data from an airborne Doppler wind lidar in preparation for validation by the airborne system of an earth orbiting Doppler wind lidar sensor.

Kavaya, Michael J.↗

Test Bed Doppler Wind Lidar and Intercomparison Facility At NASA Langley Research Center

State of the art 2-micron lasers and other lidar components under development by NASA are being demonstrated and validated in a mobile test bed Doppler wind lidar. A lidar intercomparison facility has been developed to ensure parallel alignment of up to 4 Doppler lidar systems while measuring wind. Investigations of the new components; their operation in a complete system; systematic and random errors; the hybrid (joint coherent and direct detection) approach to global wind measurement; and atmospheric wind behavior are planned. Future uses of the VALIDAR (VALIDation LIDAR) mobile lidar may include comparison with the data from an airborne Doppler wind lidar in preparation for validation by the airborne system of an earth orbiting Doppler wind lidar sensor.

Kavaya, Michael J.↗

(abstract) A High Throughput 3-D Inner Product Processor

A particularily challenging image processing application is the real time scene acquisition and object discrimination. It requires spatio-temporal recognition of point and resolved objects at high speeds with parallel processing algorithms. Neural network paradigms provide fine grain parallism and, when implemented in hardware, offer orders of magnitude speed up. However, neural networks implemented on a VLSI chip are planer architectures capable of efficient processing of linear vector signals rather than 2-D images. Therefore, for processing of images, a 3-D stack of neural-net ICs receiving planar inputs and consuming minimal power are required. Details of the circuits with chip architectures will be described with need to develop ultralow-power electronics. Further, use of the architecture in a system for high-speed processing will be illustrated.

imaging parallel processing algorithms linear vect↗