Search NASA⌕ Search

SEARCH · Search NASA

Results for “Multiple processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Waveform Developer's Guide for the Integrated Power, Avionics, and Software (iPAS) Space Telecommunications Radio System (STRS) Radio

The Space Telecommunications Radio System (STRS) provides a common, consistent framework for software defined radios (SDRs) to abstract the application software from the radio platform hardware. The STRS standard aims to reduce the cost and risk of using complex, configurable and reprogrammable radio systems across NASA missions. To promote the use of the STRS architecture for future NASA advanced exploration missions, NASA Glenn Research Center (GRC) developed an STRS-compliant SDR on a radio platform used by the Advance Exploration System program at the Johnson Space Center (JSC) in their Integrated Power, Avionics, and Software (iPAS) laboratory. The iPAS STRS Radio was implemented on the Reconfigurable, Intelligently-Adaptive Communication System (RIACS) platform, currently being used for radio development at JSC. The platform consists of a Xilinx(Trademark) ML605 Virtex(Trademark)-6 FPGA board, an Analog Devices FMCOMMS1-EBZ RF transceiver board, and an Embedded PC (Axiomtek(Trademark) eBox 620-110-FL) running the Ubuntu 12.4 operating system. The result of this development is a very low cost STRS compliant platform that can be used for waveform developments for multiple applications. The purpose of this document is to describe how to develop a new waveform using the RIACS platform and the Very High Speed Integrated Circuits (VHSIC) Hardware Description Language (VHDL) FPGA wrapper code and the STRS implementation on the Axiomtek processor.

Software Defined Radio↗

Parallel methods for dynamic simulation of multiple manipulator systems

In this paper, efficient dynamic simulation algorithms for a system of m manipulators, cooperating to manipulate a large load, are developed; their performance, using two possible forms of parallelism on a general-purpose parallel computer, is investigated. One form, temporal parallelism, is obtained with the use of parallel numerical integration methods. A speedup of 3.78 on four processors of CRAY Y-MP8 was achieved with a parallel four-point block predictor-corrector method for the simulation of a four manipulator system. These multi-point methods suffer from reduced accuracy, and when comparing these runs with a serial integration method, the speedup can be as low as 1.83 for simulations with the same accuracy. To regain the performance lost due to accuracy problems, a second form of parallelism is employed. Spatial parallelism allows most of the dynamics of each manipulator chain to be computed simultaneously. Used exclusively in the four processor case, this form of parallelism in conjunction with a serial integration method results in a speedup of 3.1 on four processors over the best serial method. In cases where there are either more processors available or fewer chains in the system, the multi-point parallel integration methods are still advantageous despite the reduced accuracy because both forms of parallelism can then combine to generate more parallel tasks and achieve greater effective speedups. This paper also includes results for these cases.

Mcmillan, Scott↗

Gain in computational efficiency by vectorization in the dynamic simulation of multi-body systems

An improved technique for the identification and extraction of the exact quantities associated with the degrees of freedom at the element as well as the flexible body level is presented. It is implemented in the dynamic equations of motions based on the recursive formulation of Kane et al. (1987) and presented in a matrix form, integrating the concepts of strain energy, the finite-element approach, modal analysis, and reduction of equations. This technique eliminates the CPU intensive matrix multiplication operations in the code's hot spots for the dynamic simulation of the interconnected rigid and flexible bodies. A study of a simple robot with flexible links is presented by comparing the execution times on a scalar machine and a vector-processor with and without vector options. Performance figures demonstrating the substantial gains achieved by the technique are plotted.

Amirouche, F. M. L.↗

An electronic pan/tilt/zoom camera system

A small camera system is described for remote viewing applications that employs fisheye optics and electronics processing for providing pan, tilt, zoom, and rotational movements. The fisheye lens is designed to give a complete hemispherical FOV with significant peripheral distortion that is corrected with high-speed electronic circuitry. Flexible control of the viewing requirements is provided by a programmable transformation processor so that pan/tilt/rotation/zoom functions can be accomplished without mechanical movements. Images are presented that were taken with a prototype system using a CCD camera, and 5 frames/sec can be acquired from a 180-deg FOV. The image-tranformation device can provide multiple images with different magnifications and pan/tilt/rotation sequences at frame rates compatible with conventional video devices. The system is of interest to the object tracking, surveillance, and viewing in constrained environments that would require the use of several cameras.

Zimmermann, Steve↗

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗

A Navier-Strokes Chimera Code on the Connection Machine CM-5: Design and Performance

We have implemented a three-dimensional compressible Navier-Stokes code on the Connection Machine CM-5. The code is set up for implicit time-stepping on single or multiple structured grids. For multiple grids and geometrically complex problems, we follow the 'chimera' approach, where flow data on one zone is interpolated onto another in the region of overlap. We will describe our design philosophy and give some timing results for the current code. A parallel machine like the CM-5 is well-suited for finite-difference methods on structured grids. The regular pattern of connections of a structured mesh maps well onto the architecture of the machine. So the first design choice, finite differences on a structured mesh, is natural. We use centered differences in space, with added artificial dissipation terms. When numerically solving the Navier-Stokes equations, there are liable to be some mesh cells near a solid body that are small in at least one direction. This mesh cell geometry can impose a very severe CFL (Courant-Friedrichs-Lewy) condition on the time step for explicit time-stepping methods. Thus, though explicit time-stepping is well-suited to the architecture of the machine, we have adopted implicit time-stepping. We have further taken the approximate factorization approach. This creates the need to solve large banded linear systems and creates the first possible barrier to an efficient algorithm. To overcome this first possible barrier we have considered two options. The first is just to solve the banded linear systems with data spread over the whole machine, using whatever fast method is available. This option is adequate for solving scalar tridiagonal systems, but for scalar pentadiagonal or block tridiagonal systems it is somewhat slower than desired. The second option is to 'transpose' the flow and geometry variables as part of the time-stepping process: Start with x-lines of data in-processor. Form explicit terms in x, then transpose so y-lines of data are in-processor. Form explicit terms in y, then transpose so z-lines are in processor. Form explicit terms in z, then solve linear systems in the z-direction. Transpose to the y-direction, then solve linear systems in the y-direction. Finally transpose to the x direction and solve linear systems in the x-direction. This strategy avoids inter-processor communication when differencing and solving linear systems, but requires a large amount of communication when doing the transposes. The transpose method is more efficient than the non-transpose strategy when dealing with scalar pentadiagonal or block tridiagonal systems. For handling geometrically complex problems the chimera strategy was adopted. For multiple zone cases we compute on each zone sequentially (using the whole parallel machine), then send the chimera interpolation data to a distributed data structure (array) laid out over the whole machine. This information transfer implies an irregular communication pattern, and is the second possible barrier to an efficient algorithm. We have implemented these ideas on the CM-5 using CMF (Connection Machine Fortran), a data parallel language which combines elements of Fortran 90 and certain extensions, and which bears a strong similarity to High Performance Fortran. We make use of the Connection Machine Scientific Software Library (CMSSL) for the linear solver and array transpose operations.

Jespersen, Dennis C.↗

Advanced communications technologies for image processing

It is essential for image analysts to have the capability to link to remote facilities as a means of accessing both data bases and high-speed processors. This can increase productivity through enhanced data access and minimization of delays. New technology is emerging to provide the high communication data rates needed in image processing. These developments include multi-user sharing of high bandwidth (60 megabits per second) Time Division Multiple Access (TDMA) satellite links, low-cost satellite ground stations, and high speed adaptive quadrature modems that allow 9600 bit per second communications over voice-grade telephone lines.

Likens, W. C.↗

Experimenting With Multiprocessor Simulator Concepts

Multiple microcomputer system used to investigate application of parallel processing to real-time simulation. With dual-base architecture, each microcomputer communicates with corresponding microcomputer on opposite bus through dual-port interface memory. Transfers of data to and from front-end processor occur on interactive information bus. Transfers of data related to simulation calculations occur on real-time-information bus. System, called the real-time multiprocessor simulator (RTMPS), is tool for developing low-cost, portable, user-friendly simulators.

Blech, Richard A.↗

ACTS Ka-Band Earth Stations: Technology, Performance, and Lessons Learned

The Advanced Communications Technology Satellite (ACTS) Project invested heavily in prototype Ka-band satellite ground terminals to conduct an experiments program with the ACTS satellite. The ACTS experiment's program proposed to validate Ka-band satellite and ground station technology. demonstrate future telecommunication services. demonstrate commercial viability and market acceptability of these new services, evaluate system networking and processing technology, and characterize Ka-band propagation effects, including development of techniques to mitigate signal fading. This paper will present a summary of the fixed ground terminals developed by the NASA Glenn Research Center and its industry partners, emphasizing the technology and performance of the terminals (Part 1) and the lessons learned throughout their six year operation including the inclined orbit phase of operations (Full Report). An overview of the Ka-band technology and components developed for the ACTS ground stations is presented. Next. the performance of the ground station technology and its evolution during the ACTS campaign are discussed to illustrate the technical tradeoffs made during the program and highlight technical advances by industry to support the ACTS experiments program and terminal operations. Finally. lessons learned during development and operation of the user terminals are discussed for consideration of commercial adoption into future Ka-band systems. The fixed ground stations used for experiments by government, academic, and commercial entities used reflector based offset-fed antenna systems ranging in size from 0.35m to 3.4m antenna diameter. Gateway earth stations included two systems, referred to as the NASA Ground Station (NGS) and the Link Evaluation Terminal (LET). The NGS provides tracking, telemetry, and control (TT&C) and Time Division Multiple Access (TDMA) network control functions. The LET supports technology verification and high data rate experiments. The ground stations successfully demonstrated many services and applications at Ka-band in three different modes of operation: circuit switched TDMA using the satellite on-board processor, satellite switched SS-TDMA applications using the on-board Microwave Switch Matrix (MSM), and conventional transponder (bent-pipe) operation. Data rates ranged from 4.8 kbps up to 622 Mbps. Experiments included: 1) low rate (4.8- 1 00's kbps) remote data acquisition and control using small earth stations, 2) moderate rate (1-45 Mbps) experiments included full duplex voice and video conferencing and both full duplex and asymmetric data rate protocol and network evaluation using mid-size ground stations, and 3) link characterization experiments and high data rate (155-622 Mbps) terrestrial and satellite interoperability application experiments conducted by a consortium of experimenters using the large transportable ground stations.

Reinhart, Richard C.↗

An Engineering Data Management System for Ipad

An overview of the capabilities and software architecture of the IPAD information processor (IPIP) is presented. IPIP is a state-of-the-art data base management system that satisfies engineering requirements not addressed by present day commercial systems. It also significantly advances a number of capabilities that are offered commercially. IPIP capabilities range from support for multiple schemas and data models to support for distributed processing, configuration control, and data inventory management. IPIP exploits semantic commonality in features offered in various forms at different user interfaces in today's commercial systems. An integrated software architecture supports all user interfaces: programming languages, interactive data manipulation, and schema languages. This approach promotes simplicity and compactness in software and permits features to be offered symmetrically across all appropriate user interfaces.

H R Johnson↗

Application of concurrent processing to structural dynamic response computations

Described are the experiences gained from solving for the dynamic response of two simple structures on an experimental Multiple Instruction Multiple Data (MIMD) computer called the finite element machine. Introduced are MIMD computing concepts, describing how the concurrent algorithmic techniques implemented and giving results for the two example problems. The results show computational speedups of up to 7.83 using eight of the finite element machine processors and indicate that significant computational speedups are possible for large order structural computations.

Ransom, J.↗

Extended testing of a general contextual classifier using the massively parallel processor - Preliminary results and test plans

Earlier encouraging test results of a contextual classifier that combines spatial and spectral information employing a general statistical approach are expanded. The earlier results were of limited meaning because they were produced from small (50-by-50 pixel) data sets. An implementation of the contextual classifier on NASA Goddard's Massively Parallel Processor (MPP) is presented; for the first time the MPP makes feasible the testing of the classifier on large data sets (a 12-hour test on a VAX-11/780 minicomputer now takes 5 minutes on the MPP). The MPP is a Single-Instruction, Multiple Data Stream computer, consisting of 16,384 bit serial microprocessors connected in a 128-by-128 mesh array with each element having data transfer connections with its four nearest neighbors so that the MPP is capable of billions of operations per second. Preliminary results are given (with more expected for the conference) and plans are mentioned for extended testing of the contextual classifier on Thematic Mapper data sets.

Tilton, J. C.↗

Vector performance analysis of three supercomputers - Cray-2, Cray Y-MP, and ETA10-Q

Results are presented of a series of experiments to study the single-processor performance of three supercomputers: Cray-2, Cray Y-MP, and ETA10-Q. The main object of this study is to determine the impact of certain architectural features on the performance of modern supercomputers. Features such as clock period, memory links, memory organization, multiple functional units, and chaining are considered. A simple performance model is used to examine the impact of these features on the performance of a set of basic operations. The results of implementing this set on these machines for three vector lengths and three memory strides are presented and compared. For unit stride operations, the Cray Y-MP outperformed the Cray-2 by as much as three times and the ETA10-Q by as much as four times for these operations. Moreover, unlike the Cray-2 and ETA10-Q, even-numbered strides do not cause a major performance degradation on the Cray Y-MP. Two numerical algorithms are also used for comparison. For three problem sizes of both algorithms, the Cray Y-MP outperformed the Cray-2 by 43 percent to 68 percent and the ETA10-Q by four to eight times.

Fatoohi, Rod A.↗

Compact Optoelectronic Compass

A compact optoelectronic sensor unit measures the apparent motion of the Sun across the sky. The data acquired by this chip are processed in an external processor to estimate the relative orientation of the axis of rotation of the Earth. Hence, the combination of this chip and the external processor finds the direction of true North relative to the chip: in other words, the combination acts as a solar compass. If the compass is further combined with a clock, then the combination can be used to establish a threeaxis inertial coordinate system. If, in addition, an auxiliary sensor measures the local vertical direction, then the resulting system can determine the geographic position. This chip and the software used in the processor are based mostly on the same design and operation as those of the unit described in Micro Sun Sensor for Spacecraft (NPO-30867) elsewhere in this issue of NASA Tech Briefs. Like the unit described in that article, this unit includes a small multiple-pinhole camera comprising a micromachined mask containing a rectangular array of microscopic pinholes mounted a short distance in front of an image detector of the active-pixel sensor (APS) type (see figure). Further as in the other unit, the digitized output of the APS in this chip is processed to compute the centroids of the pinhole Sun images on the APS. Then the direction to the Sun, relative to the compass chip, is computed from the positions of the centroids (just like a sundial). In the operation of this chip, one is interested not only in the instantaneous direction to the Sun but also in the apparent path traced out by the direction to the Sun as a result of rotation of the Earth during an observation interval (during which the Sun sensor must remain stationary with respect to the Earth). The apparent path of the Sun across the sky is projected on a sphere. The axis of rotation of the Earth lies at the center of the projected circle on the sphere surface. Hence, true North (not magnetic North), relative to the chip, can be estimated from paths of the Sun images across the APS. In a test, this solar compass has been found to yield a coarse estimate of the North (within tens of degrees) in an observation time of about ten minutes. As expected, the accuracy was found to increase with observation time: after a few hours, the estimated direction of the rotation axis becomes accurate to within a small fraction of a degree.

Christian, Carl↗

Software fault tolerance in computer operating systems

This chapter provides data and analysis of the dependability and fault tolerance for three operating systems: the Tandem/GUARDIAN fault-tolerant system, the VAX/VMS distributed system, and the IBM/MVS system. Based on measurements from these systems, basic software error characteristics are investigated. Fault tolerance in operating systems resulting from the use of process pairs and recovery routines is evaluated. Two levels of models are developed to analyze error and recovery processes inside an operating system and interactions among multiple instances of an operating system running in a distributed environment. The measurements show that the use of process pairs in Tandem systems, which was originally intended for tolerating hardware faults, allows the system to tolerate about 70% of defects in system software that result in processor failures. The loose coupling between processors which results in the backup execution (the processor state and the sequence of events occurring) being different from the original execution is a major reason for the measured software fault tolerance. The IBM/MVS system fault tolerance almost doubles when recovery routines are provided, in comparison to the case in which no recovery routines are available. However, even when recovery routines are provided, there is almost a 50% chance of system failure when critical system jobs are involved.

Iyer, Ravishankar K.↗

Parallel integer sorting with medium and fine-scale parallelism

Two new parallel integer sorting algorithms, queue-sort and barrel-sort, are presented and analyzed in detail. These algorithms do not have optimal parallel complexity, yet they show very good performance in practice. Queue-sort designed for fine-scale parallel architectures which allow the queueing of multiple messages to the same destination. Barrel-sort is designed for medium-scale parallel architectures with a high message passing overhead. The performance results from the implementation of queue-sort on a Connection Machine CM-2 and barrel-sort on a 128 processor iPSC/860 are given. The two implementations are found to be comparable in performance but not as good as a fully vectorized bucket sort on the Cray YMP.

Dagum, Leonardo↗

Single-event upset in advanced commercial power PC microprocessors

Single-event upset from heavy ions in measured for advanced commercial microprocessors, comparing upset sensitivity in registers and d-cache for several generations of devices. Multiple-bit upsets and asymmetry in registers upset cross sections are also discussed.

single-event upset advanced commercial power PC pr↗

End-to-end protocol for high-quality quantum approximate optimization algorithm parameters with few shots

The quantum approximate optimization algorithm (QAOA) is a quantum heuristic for combinatorial optimization that has been demonstrated to scale better than state-of-the-art classical solvers for some problems. For a given problem instance, QAOA performance depends crucially on the choice of the parameters. While average-case optimal parameters are available in many cases, meaningful performance gains can be obtained by fine-tuning these parameters for a given instance. This task is especially challenging, however, when the number of circuit executions (shots) is limited. In this work, we develop an end-to-end protocol that combines multiple parameter settings and fine-tuning techniques. We use large-scale numerical experiments to optimize the protocol for the shot-limited setting and observe that optimizers with the simplest internal model (linear) perform best. We implement the optimized pipeline on a trapped-ion processor using up to 32 qubits and 5 QAOA layers, and we demonstrate that the pipeline is robust to small amounts of hardware noise. To the best of our knowledge, these are the largest demonstrations of QAOA parameter fine-tuning on a trapped-ion processor in terms of two-qubit gate count.

quantum algorithms & computation↗