Search NASASearch

SEARCH · Search NASA

Results for “Task based parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A computational-grid based system for continental drainage network extraction using SRTM digital elevation models

We describe a new effort for the computation of elevation derivatives using the Shuttle Radar Topography Mission (SRTM) results. Jet Propulsion Laboratory's (JPL) SRTM has produced a near global database of highly accurate elevation data. The scope of this database enables computing precise stream drainage maps and other derivatives on Continental scales. We describe a computing architecture for this computationally very complex task based on NASA's Information Power Grid (IPG), a distributed high performance computing network based on the GLOBUS infrastructure. The SRTM data characteristics and unique problems they present are discussed. A new algorithm for organizing the conventional extraction algorithms [1] into a cooperating parallel grid is presented as an essential component to adapt to the IPG computing structure. Preliminary results are presented for a Southern California test area, established for comparing SRTM and its results against those produced using the USGS National Elevation Data (NED) model.

stream extraction

Near-Body Grid Adaption for Overset Grids

A solution adaption capability for curvilinear near-body grids has been implemented in the OVERFLOW overset grid computational fluid dynamics code. The approach follows closely that used for the Cartesian off-body grids, but inserts refined grids in the computational space of original near-body grids. Refined curvilinear grids are generated using parametric cubic interpolation, with one-sided biasing based on curvature and stretching ratio of the original grid. Sensor functions, grid marking, and solution interpolation tasks are implemented in the same fashion as for off-body grids. A goal-oriented procedure, based on largest error first, is included for controlling growth rate and maximum size of the adapted grid system. The adaption process is almost entirely parallelized using MPI, resulting in a capability suitable for viscous, moving body simulations. Two- and three-dimensional examples are presented.

Buning, Pieter G.

Fast and Accurate Greenberger-Horne-Zeilinger Encoding Using All-to-All Interactions

The 𝑁-qubit Greenberger-Horne-Zeilinger (GHZ) state is an important resource for quantum technologies. Here, we consider the task of GHZ encoding using all-to-all interactions, which prepares the GHZ state in a special case, and is furthermore useful for quantum error correction, interaction-rate enhancement, and transmitting information using power-law interactions. The naive protocol based on parallelizing CNOT gates takes O(1)-time of Hamiltonian evolution. In this work, we propose a fast protocol that achieves GHZ encoding with high accuracy. The evolution time O⁡(log 2 ⁡𝑁/𝑁) almost saturates the theoretical limit Ω⁡(log⁡𝑁/𝑁). Moreover, the final state is close to the ideal encoded one with high fidelity >1–10 −3 , up to large system sizes 𝑁 ≲ 2000. The protocol only requires a few stages of time-independent Hamiltonian evolution; the key idea is to use the data qubit as control, and to use fast spin-squeezing dynamics generated by e.g., two-axis twisting.

quantum computation

Optimal parallel evaluation of AND trees

A quantitative analysis based on both preemptive and nonpreemptive critical-path scheduling algorithms is presently conducted for the optimal degree of parallelism required in evaluating a given AND tree. The optimal degree of parallelism is found to depend on problem complexity, precedence-graph shape, and task-time distribution along each path. In addition to demonstrating the optimality of the preemptive critical-path scheduling algorithm for evaluating an arbitrary AND tree on a fixed number of processors, the possibility of efficiently ascertaining tight bounds on the number of processors for optimal processor-time efficiency is illustrated.

Wah, Benjamin W.

A scalable multidimensional fully implicit solver for Hall magnetohydrodynamics

We propose an optimally performant fully implicit algorithm for the Hall magnetohydrodynamics (HMHD) equations based on multigrid-preconditioned Jacobian-free Newton-Krylov methods. HMHD is a challenging system to solve numerically because it supports stiff fast dispersive waves. The preconditioner is formulated using an operator-split approximate block factorization (Schur complement), informed by physics insight. We use a vector-potential formulation (instead of a magnetic field one) to allow a clean segregation of the problematic $\nabla$ x $\nabla$ x operator in the electron Ohm's law subsystem. This segregation allows the formulation of an effective damped block-Jacobi smoother for multigrid. We demonstrate by analysis that our proposed block-Jacobi iteration is convergent and has the smoothing property. The resulting HMHD solver is verified linearly with wave propagation examples, and nonlinearly with the GEM challenge reconnection problem by comparison against another HMHD code. We demonstrate the excellent algorithmic and parallel performance of the algorithm up to 16384 MPI tasks in two dimensions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING

Accelerating Neutrino Event Generation in MARLEY Using CUDA-Based RNG and GPU Parallelization

MARLEY is a simulation tool that helps scientists study how low-energy neutrinos interact with matter. To work properly, MARLEY uses random numbers thousands of times in each simulation. These random numbers are important for modeling things like how neutrinos collide with atoms and what particles they produce. Right now, MARLEY runs on a regular computer processor (CPU) and uses a built-in random number generator called the Mersenne Twister. This setup works, but it can be slow, especially when trying to simulate many events. This research focuses on making MARLEY run faster by moving the random number generation and some of the repetitive calculations from the CPU to a graphics processing unit (GPU), which can handle many tasks at the same time. We use CUDA (a tool for programming NVIDIA GPUs) and cuRAND (a GPU-based random number library) to test faster alternatives to the current random number system. We compare different GPU-based generators, like curand_mtgp32, xorwow, and philox, to see which ones are the quickest and still give reliable results. Early tests show that using the GPU can make MARLEY simulations much faster. This project not only helps improve current simulation performance but also moves closer to a full simulation chain where all stages can run on modern GPU hardware.

Dunkley, Kimieka [Florida A-M]

Correlation filters for orientation estimation

An important task in many vision applications is that of rapidly estimating the orientation of an object with respect to some frame of reference. Because of their speed and parallel processing capabilities, optical correlators should prove valuable in this application. This paper considers two algorithms for object orientation estimation based on optical correlations and presents some initial simulation results.

Kumar, B. V. K. Vijaya

Effects of a psychophysiological system for adaptive automation on performance, workload, and the event-related potential P300 component

The present study examined the effects of an electroencephalographic- (EEG-) based system for adaptive automation on tracking performance and workload. In addition, event-related potentials (ERPs) to a secondary task were derived to determine whether they would provide an additional degree of workload specificity. Participants were run in an adaptive automation condition, in which the system switched between manual and automatic task modes based on the value of each individual's own EEG engagement index; a yoked control condition; or another control group, in which task mode switches followed a random pattern. Adaptive automation improved performance and resulted in lower levels of workload. Further, the P300 component of the ERP paralleled the sensitivity to task demands of the performance and subjective measures across conditions. These results indicate that it is possible to improve performance with a psychophysiological adaptive automation system and that ERPs may provide an alternative means for distinguishing among levels of cognitive task demand in such systems. Actual or potential applications of this research include improved methods for assessing operator workload and performance.

Task Performance and Analysis

Array processor architecture

A high speed parallel array data processing architecture fashioned under a computational envelope approach includes a data base memory for secondary storage of programs and data, and a plurality of memory modules interconnected to a plurality of processing modules by a connection network of the Omega gender. Programs and data are fed from the data base memory to the plurality of memory modules and from hence the programs are fed through the connection network to the array of processors (one copy of each program for each processor). Execution of the programs occur with the processors operating normally quite independently of each other in a multiprocessing fashion. For data dependent operations and other suitable operations, all processors are instructed to finish one given task or program branch before all are instructed to proceed in parallel processing fashion on the next instruction. Even when functioning in the parallel processing mode however, the processors are not locked-step but execute their own copy of the program individually unless or until another overall processor array synchronization instruction is issued.

Barnes, George H.

Construction and demonstration of a 9-string 6 DOF force reflecting joystick for telerobotics

Confrontation with difficult manipulation tasks in hostile environments such as space, has led to the development of means to transport the human's senses, skills and cognition to the remote site. The use of advanced Telerobotics to achieve this goal is examined. A novel and universal hand controller based on a fully parallel mechanical architecture is discussed. The design and implementation of this 6 DOF force reflecting joystick is shown in relationship to the general philosophy of achieving telepresence in a man-machine system.

Lindemann, Randel

Dual-task performance with ideomotor-compatible tasks: is the central processing bottleneck intact, bypassed, or shifted in locus?

The present study examined whether the central bottleneck, assumed to be primarily responsible for the psychological refractory period (PRP) effect, is intact, bypassed, or shifted in locus with ideomotor (IM)-compatible tasks. In 4 experiments, factorial combinations of IM- and non-IM-compatible tasks were used for Task 1 and Task 2. All experiments showed substantial PRP effects, with a strong dependency between Task 1 and Task 2 response times. These findings, along with model-based simulations, indicate that the processing bottleneck was not bypassed, even with two IM-compatible tasks. Nevertheless, systematic changes in the PRP and correspondence effects across experiments suggest that IM compatibility shifted the locus of the bottleneck. The findings favor an engage-bottleneck-later hypothesis, whereby parallelism between tasks occurs deeper into the processing stream for IM- than for non-IM-compatible tasks, without the bottleneck being actually eliminated.

Psychomotor Performance

Parallel Processing Systems for Passive Ranging During Helicopter Flight

The complexity of rotorcraft missions involving operations close to the ground result in high pilot workload. In order to allow a pilot time to perform mission-oriented tasks, sensor-aiding and automation of some of the guidance and control functions are highly desirable. Images from an electro-optical sensor provide a covert way of detecting objects in the flight path of a low-flying helicopter. Passive ranging consists of processing a sequence of images using techniques based on optical low computation and recursive estimation. The passive ranging algorithm has to extract obstacle information from imagery at rates varying from five to thirty or more frames per second depending on the helicopter speed. We have implemented and tested the passive ranging algorithm off-line using helicopter-collected images. However, the real-time data and computation requirements of the algorithm are beyond the capability of any off-the-shelf microprocessor or digital signal processor. This paper describes the computational requirements of the algorithm and uses parallel processing technology to meet these requirements. Various issues in the selection of a parallel processing architecture are discussed and four different computer architectures are evaluated regarding their suitability to process the algorithm in real-time. Based on this evaluation, we conclude that real-time passive ranging is a realistic goal and can be achieved with a short time.

Sridhar, Bavavar

Influence of combined visual and vestibular cues on human perception and control of horizontal rotation

Measurements are made of manual control performance in the closed-loop task of nulling perceived self-rotation velocity about an earth-vertical axis. Self-velocity estimation is modeled as a function of the simultaneous presentation of vestibular and peripheral visual field motion cues. Based on measured low-frequency operator behavior in three visual field environments, a parallel channel linear model is proposed which has separate visual and vestibular pathways summing in a complementary manner. A dual-input describing function analysis supports the complementary model; vestibular cues dominate sensation at higher frequencies. The describing function model is extended by the proposal of a nonlinear cue conflict model, in which cue weighting depends on the level of agreement between visual and vestibular cues.

Zacharias, G. L.

Test and training simulator for ground-based teleoperated in-orbit servicing

For the Post-IOC(In-Orbit Construction)-Phase of COLUMBUS it is intended to use robotic devices for the routine operations of ground-based teleoperated In-Orbit Servicing. A hardware simulator for verification of the relevant in-orbit operations technologies, the Servicing Test Facility, is necessary which mainly will support the Flight Control Center for the Manned Space-Laboratories for operational specific tasks like system simulation, training of teleoperators, parallel operation simultaneously to actual in-orbit activities and for the verification of the ground operations segment for telerobotics. The present status of definition for the facility functional and operational concept is described.

Schaefer, Bernd E.

Variable-Complexity Multidisciplinary Optimization on Parallel Computers

This report covers work conducted under grant NAG1-1562 for the NASA High Performance Computing and Communications Program (HPCCP) from December 7, 1993, to December 31, 1997. The objective of the research was to develop new multidisciplinary design optimization (MDO) techniques which exploit parallel computing to reduce the computational burden of aircraft MDO. The design of the High-Speed Civil Transport (HSCT) air-craft was selected as a test case to demonstrate the utility of our MDO methods. The three major tasks of this research grant included: development of parallel multipoint approximation methods for the aerodynamic design of the HSCT, use of parallel multipoint approximation methods for structural optimization of the HSCT, mathematical and algorithmic development including support in the integration of parallel computation for items (1) and (2). These tasks have been accomplished with the development of a response surface methodology that incorporates multi-fidelity models. For the aerodynamic design we were able to optimize with up to 20 design variables using hundreds of expensive Euler analyses together with thousands of inexpensive linear theory simulations. We have thereby demonstrated the application of CFD to a large aerodynamic design problem. For the predicting structural weight we were able to combine hundreds of structural optimizations of refined finite element models with thousands of optimizations based on coarse models. Computations have been carried out on the Intel Paragon with up to 128 nodes. The parallel computation allowed us to perform combined aerodynamic-structural optimization using state of the art models of a complex aircraft configurations.

Grossman, Bernard

A transputer based finite element solver

The feasibility of performing FEM structural-mechanics analyses on transputer systems is investigated experimentally. Transputers are programmable microprocessors equipped with local memory and point-to-point communication links; they can be joined in a large concurrent system via a programming language which supports distributed processing; this permits parallel processing at relatively low hardware cost. The computational tasks required by FEM programs are reviewed; the hardware (one PC, one master transputer, and 12 slave transputers) employed in the test calculations is described; and results demonstrating the speed and efficiency of the transputer array in assembling a global stiffness matrix and performing Gauss-Jordan matrix inversion are presented in graphs. It is predicted that larger transputer networks could approach the power of supercomputers at minicomputer costs.

Favenesi, J. A.

Optical Signature Analysis of Tumbling Rocket Bodies via Laboratory Measurements

The NASA Orbital Debris Program Office has acquired telescopic lightcurve data on massive intact objects, specifically spent rocket bodies, in order to ascertain tumble rates in support of the Active Debris Removal (ADR) task to help remediate the LEO environment. Rotation rates are needed to plan and develop proximity operations for potential future ADR operations. To better characterize and model optical data acquired from ground-based telescopes, the Optical Measurements Center (OMC) at NASA/JSC emulates illumination conditions in space using equipment and techniques that parallel telescopic observations and source-target-sensor orientations. The OMC employs a 75-watt Xenon arc lamp as a solar simulator, an SBIG CCD camera with standard Johnson/Bessel filters, and a robotic arm to simulate an object's position and rotation. The light source is mounted on a rotary arm, allowing access any phase angle between 0 -- 360 degrees. The OMC does not attempt to replicate the rotation rates, but focuses on how an object is rotating as seen from multiple phase angles. The two targets studied are scaled (1:48), SL-8 Cosmos 3M second stages. The first target is painted in the standard government "gray" scheme and the second target is primary white, as used for commercial missions. This paper summarizes results of the two scaled rocket bodies, each rotated about two primary axes: (a) a spin-stabilized rotation and (b) an end-over-end rotation. The two rotation states are being investigated as a basis for possible spin states of rocket bodies, beginning with simple spin states about the two primary axes. The data will be used to create a database of potential spin states for future works to convolve with more complex spin states. The optical signatures will be presented for specific phase angles for each rocket body and shown in conjunction with acquired optical data from multiple telescope sources.

Cowardin, H.