Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallelize algorithm computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

NASA Tech Briefs, November 1995

The contents include: 1) Mission Accomplished; 2) Resource Report: Marshall Space Flight Center; 3) NASA 1995 Software of the Year Award; 4) Microbolometers Based on Epitaxial YBa2Cu3O(sub 7-x) Thin Films; 5) Garnet Random-Access Memory; 6) Fabrication of SNS Weak Links on SOS Substrates; 7) High-Voltage MOSFET Switching Circuit; 8) Asymmetric Switching for a PWM H-Bridge Power Circuit; 9) Better Ohmic Contacts for InP Semiconductor Devices; 10) Low-Bandgap Thermovoltaic Materials and Devices; 11) Digital Frequency-Differencing Circuit; 12) Imaging Magnetometer; 13) Computer-Assisted Monitoring of a Complex System; 14) Buffered Telemetry Demodulator; 15) Compact Multifunction Inspection Head; 16) Optical Detection of Fractures in Ceramic Diaphragms; 17) Eddy-Current Detection of Cracks in Reinforced Carbon/Carbon; 18) Apparent Thermal Conductivity of Multilayer Insulation; 19) Optimizing Misch-Metal Compositions in Metal Hydride Anodes; 20) Device for Sampling Surface Contamination; 21) Probabilistic Failure Assessment for Fatigue; 22) Probabilistic Fatigue and Flaw-Propagation Analysis; 23) Windows Program for Driving the TDU-850 Printer; 24) Subband/Transform MATLAB Functions for Processing Images; 25) Computing Equilibrium Chemical Compositions; 26) Program Processes Thermocouple Readings; 27) ICAN-Second-Generation Integrated Composite Analyzer; 28) Integrated Composite Analyzer with Damping Capabilities; 29) Computing Efficiency of Transfer of Microwave Power; 30) Program Calculates Power Demands of Electronic Designs; 31) Cost-Estimation Program; 32) Program Estimates Areas Required by Electronic Designs; 33) Program to Balance Mapped Turbopump Assemblies; 34) BiblioTech; 35) Controlling Mirror Tilt With a Bimorph Actuator; 36) Burst-Disk Device Simulates Effect of Pyrotechnic Device; 37) Bearing-Mounting Concept Accommodates Thermal Expansion; 38) Parallel-Plate Acoustic Absorbers for Hot Environments; 39) Adjustable-Length Strut Withstands Large Cyclic Loads; 40) Tool Indicates Contact Angles in Bearing Raceways; 41) Gravity Slides With Magnetic Braking; 42) High-Torque, Lightweight, Pneumatically Driven Wrench for Small Spaces; 43) Device for Testing Compatibility of an O-Ring; 44) Magnetic Heat Pump Containing Flow Diverters; 45) Variable-Tilt Helicopter Rotor Mast; 46) "Beach-Ball" Robotic Rovers; 47) Apparatus Would Measure Temperatures of Ball Bearings; 48) Flexible Borescope for Inspecting Ducts; 49) Texturing Copper To Reduce Secondary Emission of Electrons; 50) Automated Laser Cutting in Three Dimensions; 51) Algorithm Helps Monitor Engine Operation; 52) Flexible Revision of Data-Processing Communications; 53) Software for Managing the Use of Land; 54) Thermal Strap Increases Cryocooling Efficiency; 55) Reversible Nut With Engagement Indication; 56) Control Algorithms for Kinematically Redundant Manipulators; 57) Computed Hydrogen-Flow Splits in a Rocket Engine; 58) Pressure and Thermal Modeling of Rocket Launches; 59) Field of View of a Spacecraft Antenna: Analysis and Software; 60) Digital Controller for Laser-Beam-Steering Subsystem; 61) More About Beam-Steering Subsystem for Laser Communication; 62) Digital Controller for Laser-Beam-Steering Subsystem: Part 2; 63) Interface Circuit Board for Space-Shuttle Communications; 64) Automated Planning of Spacecraft Telecommunications; 65) Artifacts of Spectral Analysis of Instrument Readings; 66) Neural-Network Controller for Vibration Suppression; 67) Adaptive Finite-Element Computation in Fracture Mechanics; 68) Attitude Control for the Cassini Spacecraft; 69) Analytical Model for Fluid Dynamics in a Microgravity Environment; 70) Study of Rocket-Engine Joints Bonded by NVCU/NARloy-Z; 71) Improved Silicon Nitride for Advanced Heat Engines; 72) Parameters for Welding Aluminum/Lithium Alloys; 73) Lightweight Composite Intertank Structure; 74) Foil Patches Seal Small Vacuum Leaks; 75) Data Base on Cables and Connectors; 76) Effect of Clock Mode on Radiation Hardnessf an ADC; and 77) Fault-Tolerant Control for a Robotic Inspection System.

Source record↗

The alignment-distribution graph

Implementing a data-parallel language such as Fortran 90 on a distributed-memory parallel computer requires distributing aggregate data objects (such as arrays) among the memory modules attached to the processors. The mapping of objects to the machine determines the amount of residual communication needed to bring operands of parallel operations into alignment with each other. We present a program representation called the alignment-distribution graph that makes these communication requirements explicit. We describe the details of the representation, show how to model communication cost in this framework, and outline several algorithms for determining object mappings that approximately minimize residual communication.

Chatterjee, Siddhartha↗

An Investigation of Parallel Programming Techniques Applied to Monte Carlo Simulations for Post-Flight Reconstruction of Spacecraft Trajectory

Parallelizing software to execute on multi-core central processing units (CPUs) and graphics processing units (GPUs) can be challenging. For some fields outside of Computer Science, this transition comes with new issues. For example, memory limitations can require modifications to code not initially developed to run on GPUs. This work applies the Open Multi-Processing (OpenMP) and Open Accelerators (OpenACC) directive-based parallelization strategies on a Monte Carlo simulation approach for trajectory reconstruction enabling it to run on multi-core CPUs and GPUs. Large matrix operations are the most common use of GPUs, which are not present in this algorithm; however, the natural parallelism of independent trajectories in Monte Carlo simulations is exploited. Benchmarking data are presented comparing execution times of the software for single-thread CPUs, multi-thread CPUs with OpenMP, and multi-thread GPUs using OpenACC. These data were collected using nodes with Intel® Xeon® E5-2670 (Sandy Bridge) CPUs enhanced with NVIDIA® Tesla® K40 GPUs on the Pleiades Supercomputer cluster at the National Aeronautics and Space Administration (NASA) Ames Research Center (ARC) and a local Intel® Xeon Phi™ node at NASA Langley Research Center (LaRC).

Williams, R. Anthony↗

A parallel-pipelined architecture for a multi carrier demodulator

Analog devices have been used for processing the information on board the satellites. Presently, digital devices are being used because they are economical and flexible as compared to their analog counterparts. Several schemes of digital transmission can be used depending on the data rate requirement of the user. An economical scheme of transmission for small earth stations uses single channel per carrier/frequency division multiple access (SCPC/FDMA) on the uplink and time division multiplexing (TDM) on the downlink. This is a typical communication service offered to low data rate users in commercial mass market. These channels usually pertain to either voice or data transmission. An efficient digital demodulator architecture is provided for a large number of law data rate users. A demodulator primarily consists of carrier, clock, and data recovery modules. This design uses principles of parallel processing, pipelining, and time sharing schemes to process large numbers of voice or data channels. It maintains the optimum throughput which is derived from the designed architecture and from the use of high speed components. The design is optimized for reduced power and area requirements. This is essential for satellite applications. The design is also flexible in processing a group of a varying number of channels. The algorithms that are used are verified by the use of a computer aided software engineering (CASE) tool called the Block Oriented System Simulator. The data flow, control circuitry, and interface of the hardware design is simulated in C language. Also, a multiprocessor approach is provided to map, model, and simulate the demodulation algorithms mainly from a speed view point. A hypercude based architecture implementation is provided for such a scheme of operation. The hypercube structure and the demodulation models on hypercubes are simulated in Ada.

Kwatra, S. C.↗

Parallel Computation of Unsteady Flows on a Network of Workstations

Parallel computation of unsteady flows requires significant computational resources. The utilization of a network of workstations seems an efficient solution to the problem where large problems can be treated at a reasonable cost. This approach requires the solution of several problems: 1) the partitioning and distribution of the problem over a network of workstation, 2) efficient communication tools, 3) managing the system efficiently for a given problem. Of course, there is the question of the efficiency of any given numerical algorithm to such a computing system. NPARC code was chosen as a sample for the application. For the explicit version of the NPARC code both two- and three-dimensional problems were studied. Again both steady and unsteady problems were investigated. The issues studied as a part of the research program were: 1) how to distribute the data between the workstations, 2) how to compute and how to communicate at each node efficiently, 3) how to balance the load distribution. In the following, a summary of these activities is presented. Details of the work have been presented and published as referenced.

Source record↗

Reconfigurable Hardware for Compressing Hyperspectral Image Data

High-speed, low-power, reconfigurable electronic hardware has been developed to implement ICER-3D, an algorithm for compressing hyperspectral-image data. The algorithm and parts thereof have been the topics of several NASA Tech Briefs articles, including Context Modeler for Wavelet Compression of Hyperspectral Images (NPO-43239) and ICER-3D Hyperspectral Image Compression Software (NPO-43238), which appear elsewhere in this issue of NASA Tech Briefs. As described in more detail in those articles, the algorithm includes three main subalgorithms: one for computing wavelet transforms, one for context modeling, and one for entropy encoding. For the purpose of designing the hardware, these subalgorithms are treated as modules to be implemented efficiently in field-programmable gate arrays (FPGAs). The design takes advantage of industry- standard, commercially available FPGAs. The implementation targets the Xilinx Virtex II pro architecture, which has embedded PowerPC processor cores with flexible on-chip bus architecture. It incorporates an efficient parallel and pipelined architecture to compress the three-dimensional image data. The design provides for internal buffering to minimize intensive input/output operations while making efficient use of offchip memory. The design is scalable in that the subalgorithms are implemented as independent hardware modules that can be combined in parallel to increase throughput. The on-chip processor manages the overall operation of the compression system, including execution of the top-level control functions as well as scheduling, initiating, and monitoring processes. The design prototype has been demonstrated to be capable of compressing hyperspectral data at a rate of 4.5 megasamples per second at a conservative clock frequency of 50 MHz, with a potential for substantially greater throughput at a higher clock frequency. The power consumption of the prototype is less than 6.5 W. The reconfigurability (by means of reprogramming) of the FPGAs makes it possible to effectively alter the design to some extent to satisfy different requirements without adding hardware. The implementation could be easily propagated to future FPGA generations and/or to custom application-specific integrated circuits.

Aranki, Nazeeh↗

Temporal Planning for Compilation of Quantum Approximate Optimization Algorithm Circuits

We investigate the application of temporal planners to the problem of compiling quantum circuits to newly emerging quantum hardware. While our approach is general, we focus our initial experiments on Quantum Approximate Optimization Algorithm (QAOA) circuits that have few ordering constraints and allow highly parallel plans. We report on experiments using several temporal planners to compile circuits of various sizes to a realistic hardware. This early empirical evaluation suggests that temporal planning is a viable approach to quantum circuit compilation.

planning↗

DFT algorithms for bit-serial GaAs array processor architectures

Systems and Processes Engineering Corporation (SPEC) has developed an innovative array processor architecture for computing Fourier transforms and other commonly used signal processing algorithms. This architecture is designed to extract the highest possible array performance from state-of-the-art GaAs technology. SPEC's architectural design includes a high performance RISC processor implemented in GaAs, along with a Floating Point Coprocessor and a unique Array Communications Coprocessor, also implemented in GaAs technology. Together, these data processors represent the latest in technology, both from an architectural and implementation viewpoint. SPEC has examined numerous algorithms and parallel processing architectures to determine the optimum array processor architecture. SPEC has developed an array processor architecture with integral communications ability to provide maximum node connectivity. The Array Communications Coprocessor embeds communications operations directly in the core of the processor architecture. A Floating Point Coprocessor architecture has been defined that utilizes Bit-Serial arithmetic units, operating at very high frequency, to perform floating point operations. These Bit-Serial devices reduce the device integration level and complexity to a level compatible with state-of-the-art GaAs device technology.

Mcmillan, Gary B.↗

Adaptive pattern recognition by mini-max neural networks as a part of an intelligent processor

In this decade and progressing into 21st Century, NASA will have missions including Space Station and the Earth related Planet Sciences. To support these missions, a high degree of sophistication in machine automation and an increasing amount of data processing throughput rate are necessary. Meeting these challenges requires intelligent machines, designed to support the necessary automations in a remote space and hazardous environment. There are two approaches to designing these intelligent machines. One of these is the knowledge-based expert system approach, namely AI. The other is a non-rule approach based on parallel and distributed computing for adaptive fault-tolerances, namely Neural or Natural Intelligence (NI). The union of AI and NI is the solution to the problem stated above. The NI segment of this unit extracts features automatically by applying Cauchy simulated annealing to a mini-max cost energy function. The feature discovered by NI can then be passed to the AI system for future processing, and vice versa. This passing increases reliability, for AI can follow the NI formulated algorithm exactly, and can provide the context knowledge base as the constraints of neurocomputing. The mini-max cost function that solves the unknown feature can furthermore give us a top-down architectural design of neural networks by means of Taylor series expansion of the cost function. A typical mini-max cost function consists of the sample variance of each class in the numerator, and separation of the center of each class in the denominator. Thus, when the total cost energy is minimized, the conflicting goals of intraclass clustering and interclass segregation are achieved simultaneously.

Szu, Harold H.↗

Mission-Maps For Outbound Cislunar Transfer Trajectories

This study quantifies the robustness and sensitivity of an outbound cislunar trajectory for a lunar lander in the form of mission-maps, or topological maps that allows either a computer program or mission designer to intuitively optimize the placement of critical outbound correction burns from the derived sensitivity data. The non-linear multi-body dynamics are applied to generate an outbound cislunar reference profile used by a linear covariance analysis (LinCov) tool to compute the expected Δv and trajectory dispersions due to the initial state uncertainty, sensor errors, maneuver execution errors, and disturbance accelerations along the outbound cislunar profile. The rapid performance analysis capabilities of LinCov are complimented with parallel processing techniques to evaluate hundreds and thousands of different translational burn locations, placements, and targeting constraints to identify the combination that minimizes the total Δv usage (nominal plus 3σ Δv) and trajectory dispersions at lunar orbit insertion. This study utilizes a generalized reference targeting algorithm to quickly assess the integrated closed-loop GN&C system performance due to different targeting configurations and constraints. The resulting mission maps provide an intuitive insight to ascertain each trajectory correction maneuver’s (TCM) sensitivity to different burn times along an outbound cislunar trajectory and quickly identify desirable engineering tradeoffs when performing analysis on the number and placement of these burns that nominally zero. Multiple mission maps are generated for a variety of different performance parameters that allow engineers to visually identify optimal solutions for trajectory correction maneuver placements, the number of correction burns, and the targeting constraints for each burn.

GN&C↗

High performance architecture for robot control

Practical aspects of the design and implementation of a modular, high performance, parallel computer control system for telerobots are discussed. Topics of consideration include system architecture, operator interface, and control execution. In a laboratory environment, a telerobotics test control configuration is used to obtain measurements on communications and control loop timing for use in an effective full scale operational system design. The feasibility of the selected architectural approach has been successfully demonstrated. The modularity of the software and hardware enables ease of transport for use in the operational system. The distributed partioning of the control algorithms and the performance measurements acquired during control system implementation are discussed.

Byler, E.↗

Reducing Speckle In One-Look SAR Images

Local-adaptive-filter algorithm incorporated into digital processing of synthetic-aperture-radar (SAR) echo data to reduce speckle in resulting imagery. Involves use of image statistics in vicinity of each picture element, in conjunction with original intensity of element, to estimate brightness more nearly proportional to true radar reflectance of corresponding target. Increases ratio of signal to speckle noise without substantial degradation of resolution common to multilook SAR images. Adapts to local variations of statistics within scene, preserving subtle details. Computationally simple. Lends itself to parallel processing of different segments of image, making possible increased throughput.

Nathan, K. S.↗

Adaptive Strategies for Controls of Flexible Arms

An adaptive controller for a modern manipulator has been designed based on asymptotical stability via the Lyapunov criterion with the output error between the system and a reference model used as the actuating control signal. Computer simulations were carried out to test the design. The combination of the adaptive controller and a system vibration and mode shape estimator show that the flexible arm should move along a pre-defined trajectory with high-speed motion and fast vibration setting time. An existing computer-controlled prototype two link manipulator, RALF (Robotic Arm, Large Flexible), with a parallel mechanism driven by hydraulic actuators was used to verify the mathematical analysis. The experimental results illustrate that assumed modes found from finite element techniques can be used to derive the equations of motion with acceptable accuracy. The robust adaptive (modal) control is implemented to compensate for unmodelled modes and nonlinearities and is compared with the joint feedback control in additional experiments. Preliminary results show promise for the experimental control algorithm.

Yuan, Bau-San↗

Lunar rovers and local positioning system

Telerobotic rovers equipped with adequate actuators and sensors are clearly necessary for extraterrestrial construction. They will be employed as substitutes for humans, to perform jobs like surveying, sensing, signaling, manipulating, and the handling of small materials. Important design criteria for these rovers include versatility and robustness. They must be easily programmed and reprogrammed to perform a wide variety of different functions, and they must be robust so that construction work will not be jeopardized by parts failures. The key qualities and functions necessary for these rovers to achieve the required versatility and robustness are modularity, redundancy, and coordination. Three robotic rovers are being built by CSC as a test bed to implement the concepts of modularity and coordination. The specific goal of the design and construction of these robots is to demonstrate the software modularity and multirobot control algorithms required for the physical manipulation of constructible elements. Each rover consists of a transporter platform, bus manager, simple manipulator, and positioning receivers. These robots will be controlled from a central control console via a radio-frequency local area network (LAN). To date, one prototype transporter platform frame was built with batteries, motors, a prototype single-motor controller, and two prototype internal LAN boards. Software modules were developed in C language for monitor functions, i/o, and parallel port usage in each computer board. Also completed are the fabrication of half of the required number of computer boards, the procurement of 19.2 Kbaud RF modems for inter-robot communications, and the simulation of processing requirements for positioning receivers. In addition to the robotic platform, the fabrication of a local positioning system based on infrared signals is nearly completed. This positioning system will make the rovers into a moving reference system capable of performing site surveys. In addition, a four degree mechanical manipulator especially suited for coordinated teleoperation was conceptually designed and is currently being analyzed. This manipulator will be integrated into the rovers as their end effector. Twenty internal LAN cards fabricated by a commercial firm are being used, a prototype manipulator and a range finder for a positioning system were built, a prototype two-motor controller was designed, and one of the robots is performing its first telerobotic motion. In addition, the robots' internal LAN's were coordinated and tested, hardware design upgrades based on fabrication and fit experience were completed, and the positioning system is running. The rover system is able to perform simple tasks such as sensing and signaling; coordination systems which allow construction tasks to begin were established, and soon coordinated teams of robots in the laboratory will be able to manipulate common objects.

Avery, James↗

Feed-forward volume rendering algorithm for moderately parallel MIMD machines

Algorithms for direct volume rendering on parallel and vector processors are investigated. Volumes are transformed efficiently on parallel processors by dividing the data into slices and beams of voxels. Equal sized sets of slices along one axis are distributed to processors. Parallelism is achieved at two levels. Because each slice can be transformed independently of others, processors transform their assigned slices with no communication, thus providing maximum possible parallelism at the first level. Within each slice, consecutive beams are incrementally transformed using coherency in the transformation computation. Also, coherency across slices can be exploited to further enhance performance. This coherency yields the second level of parallelism through the use of the vector processing or pipelining. Other ongoing efforts include investigations into image reconstruction techniques, load balancing strategies, and improving performance.

Yagel, Roni↗

New Estimates of Hydrological and Oceanic Excitations of Variations of Earth's Rotation, Geocenter and Gravitational Field

Hydrological mass transport in the geophysical fluids of the atmosphere-hydrosphere-solid Earth surface system can excite Earth's rotational variations in both length-of-day and polar motion. These effects can be computed in terms of the hydrological angular momentum by proper integration of global meteorological data. We do so using the 40-year NCEP data and the 18-year NASA GEOS-1 data, where the precipitation and evapotranspiration budgets are computed via the water mass balance of the atmosphere based on Oki et al.'s (1995) algorithm. This hydrological mass redistribution will also cause geocenter motion and changes in Earth's gravitational field, which are similarly computed using the same data sets. Corresponding geodynamic effects due to the oceanic mass transports (i.e. oceanic angular momentum and ocean-induced geocenter/gravity changes) have also been computed in a similar manner. We here compare two independent sets of the result from: (1) non-steric ocean surface topography observations based on Topex/Poseidon, and (2) the model output of the mass field by the Parallel Ocean Climate Model. Finally, the hydrological and the oceanic time series are combined in an effort to better explain the observed non-atmospheric effects. The latter are obtained by subtracting the atmospheric angular momentum from Earth rotation observations, and the atmosphere- induced geocenter/gravity effects from corresponding geodetic observations, both using the above-mentioned atmospheric data sets.

Chao, Benjamin F.↗

Proceedings of the Peteflops-Systems Operation Working Review (POWR)

This report constitutes the final technical report. Even as Petaflops performance computing is being achieved for a few important applications on the nation's largest massively parallel processing (MPP) systems, the challenges of realizing far greater performance are being investigated by a strong interdisciplinary team of experts from academia, industry, and government. This team is also exploring the extraordinary opportunity such capabilities would provide for critical areas of strategic national importance. Under what has become known informally as the "Petaflops Initiative," key leaders in research across the areas of device technology, parallel systems architecture, applications and algorithms, and systems software have been exploring the implications, requirements, interrelationships, and trade-offs among these research areas through a series of workshops, studies, and projects sponsored by a number of Federal agencies.

Sterling, Thomas↗