Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59

Delta-Doped Back-Illuminated CMOS Imaging Arrays: Progress and Prospects

In this paper, we report the latest results on our development of delta-doped, thinned, back-illuminated CMOS imaging arrays. As with charge-coupled devices, thinning and back-illumination are essential to the development of high performance CMOS imaging arrays. Problems with back surface passivation have emerged as critical to the prospects for incorporating CMOS imaging arrays into high performance scientific instruments, just as they did for CCDs over twenty years ago. In the early 1990's, JPL developed delta-doped CCDs, in which low temperature molecular beam epitaxy was used to form an ideal passivation layer on the silicon back surface. Comprising only a few nanometers of highly-doped epitaxial silicon, delta-doping achieves the stability and uniformity that are essential for high performance imaging and spectroscopy. Delta-doped CCDs were shown to have high, stable, and uniform quantum efficiency across the entire spectral range from the extreme ultraviolet through the near infrared. JPL has recently bump-bonded thinned, delta-doped CMOS imaging arrays to a CMOS readout, and demonstrated imaging. Delta-doped CMOS devices exhibit the high quantum efficiency that has become the standard for scientific-grade CCDs. Together with new circuit designs for low-noise readout currently under development, delta-doping expands the potential scientific applications of CMOS imaging arrays, and brings within reach important new capabilities, such as fast, high-sensitivity imaging with parallel readout and real-time signal processing. It remains to demonstrate manufacturability of delta-doped CMOS imaging arrays. To that end, JPL has acquired a new silicon MBE and ancillary equipment for delta-doping wafers up to 200mm in diameter, and is now developing processes for high-throughput, high yield delta-doping of fully-processed wafers with CCD and CMOS imaging devices.

imagers↗

The First Deep Space Cubesat Broadband IR Spectrometer, Lunarcubes, and the Search for Lunar Volatiles

BIRCHES is the compact broadband IR spectrometer of the Lunar Ice Cube mission. Lunar Ice Cube is one of 13 6U cubesats that will be deployed by EM1 in cislunar space, qualifying as lunarcubes. The LunarCube paradigm is a proposed approach for extending the affordable CubeSat standard to support access to deep space via cis-lunar/lunar missions. Because the lunar environment contains analogs of most solar system environments, the Moon is an ideal target for both testing critical deep space capabilities and understanding solar system formation and processes. Effectively, as developments are occurring in parallel, 13 prototype deep space cubesats are being flown for EM1. One useful outcome of this 'experiment' will be to determine to what extent it is possible to develop a lunarcube 'bus' with standardized interfaces to all subsystems using reasonable protocols for a variety of payloads. The lunar ice cube mission was developed as the test case in a GSFC R&D study to determine whether the cubesat paradigm could be applied to deep space, science requirements driven missions, and BIRCHES was its payload. JPL's Lunar Flashlight, and Arizona State University's LunaH-Map, both also EM1 lunar orbiters, will also be deployed from EM1 and provide complimentary observations to be used in understanding volatile dynamics in the same time frame.

Cubesat↗

BIRCHES and Lunarcubes: Building the First Deep Space Cubesat Broadband IR Spectrometer

The Broadband InfraRed Compact High-resolution Exploration Spectrometer (BIRCHES), which will be described in detail here, is the compact broadband IR spectrometer of the Lunar Ice Cube mission. Lunar Ice Cube is one of 13 6U cubesats that will be deployed by EM1 in cislunar space, qualifying as lunarcubes. The LunarCube paradigm is a proposed approach for extending the affordable CubeSat standard to support access to deep space via cis-lunar/lunar missions. Because the lunar environment contains analogs of most solar system environments, the Moon is an ideal target for both testing critical deep space capabilities and understanding solar system formation and processes. Effectively, as developments are occurring in parallel, 13 prototype deep space cubesats are being flown for EM1. One useful outcome of this ‘experiment’ will be to determine to what extent it is possible to develop a lunarcube ‘bus’ with standardized interfaces to all subsystems using reasonable protocols for a variety of payloads. The lunar ice cube mission was developed as the test case in a GSFC R&D study to determine whether the cubesat paradigm could be applied to deep space, science requirements driven missions, and BIRCHES was its payload. Here, we present the design and describe the ongoing development, and testing, in the context of the challenges of using the cubesat paradigm to fly a broadband IR spectrometer in a 6U platform, including minimal funding and extensive need for leveraging existing assets and relationships on development, the foreshortened schedule for payload delivery on testing, and minimum bandwidth translating into simplified or canned operation.

Chapin, Peter↗

Acceleration of Stereo Correlation in Verilog

To speed up vision processing in low speed, low power devices, embedding FPGA hardware is becoming an effective way to add processing capability. FPGAs offer the ability to flexibly add parallel and/or deeply pipelined computation to embedded processors without adding significantly to the mass and power requirements of an embedded system. This paper will discuss the JPL stereo vision system, and describe how a portion of that system was accelerated by using custom FPGA hardware to process the computationally intensive portions of JPL stereo. The architecture described takes full advantage of the ability of an FPGA to use many small computation elements in parallel. This resulted in a 16 times speedup in real hardware over using a simple linear processor to compute image correlation and disparity.

machine vision↗

On the High- and Low- Altitude Limits of the Auroral Electric Field Region

Using measurements from the High Altitude Plasma Instrument (HAPI) on the Dynamics-Explorer 1 (DE-1) spacecraft and the Low Altitude Plasma Instrument (LAPI) on Dynamics Explorer 2 (DE 2), we investigate both die high altitude and low altitude extents of the auroral acceleration region. To infer the high altitude limit, we searched the HAPI data base for evidence of upward-directed auroral electric fields located above the spacecraft when the HAPI spacecraft is above 9000 km altitude. We find that such acceleration is common when DE-1 flies through die auroral oval at an altitude of 9,000-11,000 km. At altitudes above 11,000 km, the fraction of the orbits with evidence of at least a 1000 V potential drop above the spacecraft falls, becoming essentially zero above an altitude of 15,000 km. Above that altitude, small (100 V) potential drops are frequently observed, but only rarely are approx. 1 kV potentials observed, typically associated with polar cap or 'theta' arcs or westward traveling surges. To investigate the low-altitude limit of the auroral acceleration region, we use conjunctions of DE 1 and DE 2 along auroral field lines and match the upgoing fluxes of ionospheric ions observed by DE 2 with the flux of accelerated upgoing ions observed at DE 1. Calculating the ionospheric scale height from the ion and electron temperatures and assuming that the parallel flow velocity is independent of height above 800 km, we calculate the altitude at which the upwelling ionospheric ions are effectively completely lost to upward acceleration. The initial lowest-altitude acceleration process could be either a perpendicular acceleration or a parallel electric field, but it must be sufficient to give the entire distribution escape energy. We find that in the two cases studied, near the region of peak auroral potential drop the altitude of this acceleration was around 1700 km (near the O/H neutral crossover altitude), but was significantly higher (approx. 2000 km) near the edges of the arc, where the potential was lower. The composition of the upgoing ion beam was consistent with these heights, being predominately H(+) near the edges and O(+) near the peak.

Reiff, P. H.↗

Ray Tracing Techniques for the Characterization of Lunar Communication Architectures

This paper provides an overview of the computational techniques used to characterize the viability of different lunar architectures and their ability to provide communication services to the lunar surface. This analysis was done with modern ray tracing techniques that allow for the computations to be done on Graphics Processing Unit (GPU) clusters for a high level of parallelism and severe reduction in computation time. The ray tracing computations were done with the GPU platform Compute Unified Device Architecture (CUDA) provided by NVIDIA which utilizes general-purpose computing on graphics processing units (GPGPU). This new method provides the advantage of being able to characterize a much larger portion of the lunar surface due to its computational efficiency as well as providing a more accurate representation of elevation angle limits instead of the typical and often inaccurate elevation angle mask. The Lunar surface can now be characterized with metrics such as contact time, outage time, and received data rate. With these metrics, different proposed Lunar architectures can be rapidly evaluated. This reduction in computation time not only leads to more accurate results but allows these results to be obtained in a time frame that allows for the complete characterization of the trade space. It is expected that these different architecture comparisons will lead to a conclusive determination of the optimal Lunar architecture and will allow for future Lunar missions to operate as close to real time as possible. In addition, this computation method can be used to recreate visibility figures generated by previous methods but with an increased level of accuracy.

Thomas Montano↗

Progress in the Simulation of Steady and Time-Dependent Flows with 3D Parallel Unstructured Cartesian Methods

The proposed paper will present recent extensions in the development of an efficient Euler solver for adaptively-refined Cartesian meshes with embedded boundaries. The paper will focus on extensions of the basic method to include solution adaptation, time-dependent flow simulation, and arbitrary rigid domain motion. The parallel multilevel method makes use of on-the-fly parallel domain decomposition to achieve extremely good scalability on large numbers of processors, and is coupled with an automatic coarse mesh generation algorithm for efficient processing by a multigrid smoother. Numerical results are presented demonstrating parallel speed-ups of up to 435 on 512 processors. Solution-based adaptation may be keyed off truncation error estimates using tau-extrapolation or a variety of feature detection based refinement parameters. The multigrid method is extended to for time-dependent flows through the use of a dual-time approach. The extension to rigid domain motion uses an Arbitrary Lagrangian-Eulerlarian (ALE) formulation, and results will be presented for a variety of two- and three-dimensional example problems with both simple and complex geometry.

Aftosmis, M. J.↗

Grundy - Parallel processor architecture makes programming easy

The hardware, software, and firmware of the parallel processor, Grundy, are examined. The Grundy processor uses a simple processor that has a totally orthogonal three-address instruction set. The system contains a relative and indirect processing mode to support the high-level language, and uses pseudoprocessors and read-only memory. The system supports high-level language in which arbitrary degrees of algorithmic parallelism is expressed. The functions of the compiler and invocation frame are described. Grundy uses an operating system that can be accessed by an arbitrary number of processes simultaneously, and the access time grows only as the logarithm of the number of active processes. Applications for the parallel processor are discussed.

Meier, R. J., Jr.↗

SAR processing on the MPP

The processing of synthetic aperture radar (SAR) signals using the massively parallel processor (MPP) is discussed. The fast Fourier transform convolution procedures employed in the algorithms are described. The MPP architecture comprises an array unit (ARU) which processes arrays of data; an array control unit which controls the operation of the ARU and performs scalar arithmetic; a program and data management unit which controls the flow of data; and a unique staging memory (SM) which buffers and permutes data. The ARU contains a 128 by 128 array of bit-serial processing elements (PE). Two-by-four surarrays of PE's are packaged in a custom VLSI HCMOS chip. The staging memory is a large multidimensional-access memory which buffers and permutes data flowing with the system. Efficient SAR processing is achieved via ARU communication paths and SM data manipulation. Real time processing capability can be realized via a multiple ARU, multiple SM configuration.

Batcher, K. E.↗

A design for an intelligent monitor and controller for space station electrical power using parallel distributed problem solving

The emphasis is on defining a set of communicating processes for intelligent spacecraft secondary power distribution and control. The computer hardware and software implementation platform for this work is that of the ADEPTS project at the Johnson Space Center (JSC). The electrical power system design which was used as the basis for this research is that of Space Station Freedom, although the functionality of the processes defined here generalize to any permanent manned space power control application. First, the Space Station Electrical Power Subsystem (EPS) hardware to be monitored is described, followed by a set of scenarios describing typical monitor and control activity. Then, the parallel distributed problem solving approach to knowledge engineering is introduced. There follows a two-step presentation of the intelligent software design for secondary power control. The first step decomposes the problem of monitoring and control into three primary functions. Each of the primary functions is described in detail. Suggestions for refinements and embelishments in design specifications are given.

Morris, Robert A.↗

The force on the flex: Global parallelism and portability

A parallel programming methodology, called the force, supports the construction of programs to be executed in parallel by an unspecified, but potentially large, number of processes. The methodology was originally developed on a pipelined, shared memory multiprocessor, the Denelcor HEP, and embodies the primitive operations of the force in a set of macros which expand into multiprocessor Fortran code. A small set of primitives is sufficient to write large parallel programs, and the system has been used to produce 10,000 line programs in computational fluid dynamics. The level of complexity of the force primitives is intermediate. It is high enough to mask detailed architectural differences between multiprocessors but low enough to give the user control over performance. The system is being ported to a medium scale multiprocessor, the Flex/32, which is a 20 processor system with a mixture of shared and local memory. Memory organization and the type of processor synchronization supported by the hardware on the two machines lead to some differences in efficient implementations of the force primitives, but the user interface remains the same. An initial implementation was done by retargeting the macros to Flexible Computer Corporation's ConCurrent C language. Subsequently, the macros were caused to directly produce the system calls which form the basis for ConCurrent C. The implementation of the Fortran based system is in step with Flexible Computer Corporations's implementation of a Fortran system in the parallel environment.

Jordan, H. F.↗

VLSI neuroprocessors

Electronic and optoelectronic hardware implementations of highly parallel computing architectures address several ill-defined and/or computation-intensive problems not easily solved by conventional computing techniques. The concurrent processing architectures developed are derived from a variety of advanced computing paradigms including neural network models, fuzzy logic, and cellular automata. Hardware implementation technologies range from state-of-the-art digital/analog custom-VLSI to advanced optoelectronic devices such as computer-generated holograms and e-beam fabricated Dammann gratings. JPL's concurrent processing devices group has developed a broad technology base in hardware implementable parallel algorithms, low-power and high-speed VLSI designs and building block VLSI chips, leading to application-specific high-performance embeddable processors. Application areas include high throughput map-data classification using feedforward neural networks, terrain based tactical movement planner using cellular automata, resource optimization (weapon-target assignment) using a multidimensional feedback network with lateral inhibition, and classification of rocks using an inner-product scheme on thematic mapper data. In addition to addressing specific functional needs of DOD and NASA, the JPL-developed concurrent processing device technology is also being customized for a variety of commercial applications (in collaboration with industrial partners), and is being transferred to U.S. industries. This viewgraph p resentation focuses on two application-specific processors which solve the computation intensive tasks of resource allocation (weapon-target assignment) and terrain based tactical movement planning using two extremely different topologies. Resource allocation is implemented as an asynchronous analog competitive assignment architecture inspired by the Hopfield network. Hardware realization leads to a two to four order of magnitude speed-up over conventional techniques and enables multiple assignments, (many to many), not achievable with standard statistical approaches. Tactical movement planning (finding the best path from A to B) is accomplished with a digital two-dimensional concurrent processor array. By exploiting the natural parallel decomposition of the problem in silicon, a four order of magnitude speed-up over optimized software approaches has been demonstrated.

Kemeny, Sabrina E.↗

The contribution of activated processes to Q

The possible role of activated processes in seismic attenuation is investigated. In this study, a solid is modeled by a parallel and series configuration of dashpots and springs. The contribution of stress and temperature activated processes to the long term dissipative behavior of this system is analyzed. Data from brittle rock deformation experiments suggest that one such process, stress corrosion cracking, may make a significant contribution to the attenuation factor, Q, especially for long period oscillations under significant tectonic stress.

Spetzler, H. A.↗

Head-up displays: Effect of information location on the processing of superimposed symbology

Head-up display (HUD) symbology superimposes vehicle status information onto the external terrain, providing simultaneous visual access to both sources of information. Relative to a baseline condition in which the superimposed altitude indicator was omitted, altitude maintenance was improved by the presence of the altitude indicator, and this improvement was the same magnitude regardless of the position of the altitude indicator on the screen. However, a concurrent decifit in heading maintenance was observed only when the altitude indicator was proximal to the path information. These results did not support a model of the concurrent processing deficit based on an inability to attend to multiple locations in parallel. They are consistent with previous claims that the deficit is the product of attentional limits on subjects' ability to process two separate objects (HUD symbology and terrain information) concurrently. The absence of a performance tradeoff when the HUD and the path information were less proximal is attributed to a breaking of attentional tunneling on the HUD, possibly due to eye movements.

Sanford, Beverly D.↗

Progress in Voltage and Current Mode On-Chip Analog-to-Digital converters for CMOS Image Sensors

Two 8 bit successive approximation analog-to-digital converter (ADC) designs and a 12 bit current mode incremental sigma delta ADC have been designed, fabricated, and tested. The successive approximation test chip designs are compatible with active pixel sensor (APS) column parallel architectures with a 20.4 ??itch in a 1.2 ??-well CMOS process and a 40 ??itch in a 2 ??-well CMOS process. The successive approximation designs consume as little as 49 ??t a 500 KHz conversion rate meeting the low power requirements inherent in column parallel architectures. The current mode incremental sigma-delta ADC test chips designed to be multiplied among 8 columns in a semi-column parallel current mode APS architecture. The higher accuracy ADC consumes 800 ??t a 5 KHz.

analog-to-digital↗

Aerodynamic Shape Optimization of Supersonic Aircraft Configurations via an Adjoint Formulation on Parallel Computers

This work describes the application of a control theory-based aerodynamic shape optimization method to the problem of supersonic aircraft design. The design process is greatly accelerated through the use of both control theory and a parallel implementation on distributed memory computers. Control theory is employed to derive the adjoint differential equations whose solution allows for the evaluation of design gradient information at a fraction of the computational cost required by previous design methods. The resulting problem is then implemented on parallel distributed memory architectures using a domain decomposition approach, an optimized communication schedule, and the MPI (Message Passing Interface) Standard for portability and efficiency. The final result achieves very rapid aerodynamic design based on higher order computational fluid dynamics methods (CFD). In our earlier studies, the serial implementation of this design method was shown to be effective for the optimization of airfoils, wings, wing-bodies, and complex aircraft configurations using both the potential equation and the Euler equations. In our most recent paper, the Euler method was extended to treat complete aircraft configurations via a new multiblock implementation. Furthermore, during the same conference, we also presented preliminary results demonstrating that this basic methodology could be ported to distributed memory parallel computing architectures. In this paper, our concern will be to demonstrate that the combined power of these new technologies can be used routinely in an industrial design environment by applying it to the case study of the design of typical supersonic transport configurations. A particular difficulty of this test case is posed by the propulsion/airframe integration.

Reuther, James↗

Aerodynamic Shape Optimization of Supersonic Aircraft Configurations via an Adjoint Formulation on Parallel Computers

This work describes the application of a control theory-based aerodynamic shape optimization method to the problem of supersonic aircraft design. The design process is greatly accelerated through the use of both control theory and a parallel implementation on distributed memory computers. Control theory is employed to derive the adjoint differential equations whose solution allows for the evaluation of design gradient information at a fraction of the computational cost required by previous design methods (13, 12, 44, 38). The resulting problem is then implemented on parallel distributed memory architectures using a domain decomposition approach, an optimized communication schedule, and the MPI (Message Passing Interface) Standard for portability and efficiency. The final result achieves very rapid aerodynamic design based on higher order computational fluid dynamics methods (CFD). In our earlier studies, the serial implementation of this design method (19, 20, 21, 23, 39, 25, 40, 41, 42, 43, 9) was shown to be effective for the optimization of airfoils, wings, wing-bodies, and complex aircraft configurations using both the potential equation and the Euler equations (39, 25). In our most recent paper, the Euler method was extended to treat complete aircraft configurations via a new multiblock implementation. Furthermore, during the same conference, we also presented preliminary results demonstrating that the basic methodology could be ported to distributed memory parallel computing architectures [241. In this paper, our concem will be to demonstrate that the combined power of these new technologies can be used routinely in an industrial design environment by applying it to the case study of the design of typical supersonic transport configurations. A particular difficulty of this test case is posed by the propulsion/airframe integration.

Reuther, James↗