Search NASA⌕ Search

SEARCH · Search NASA

Results for “Vectorized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

An assessment of 'shuffle algorithm' collision mechanics for particle simulations

Among the algorithms for collision mechanics used at present, the 'shuffle algorithm' of Baganoff (McDonald and Baganoff, 1988; Baganoff and McDonald, 1990) not only allows efficient vectorization, but also discretizes the possible outcomes of a collision. To assess the applicability of the shuffle algorithm, a simulation was performed of flows in monoatomic gases and the calculated characteristics of shock waves was compared with those obtained using a commonly employed isotropic scattering law. It is shown that, in general, the shuffle algorithm adequately represents the collision mechanics in cases when the goal of calculations are mean profiles of density and temperature.

Feiereisen, William J.↗

Use of the NLPQLP Sequential Quadratic Programming Algorithm to Solve Rotorcraft Aeromechanical Constrained Optimisation Problems

Optimization of the control vector, configuration and aerodynamic surface design potentially offers significant performance enhancement to rotorcraft systems. These analyses indicated that non-linear programming methods that solve a sequence of related quadratic-programming sub-problems could be used successfully to solve these problems. Accordingly, a license for one of the latest versions of Professor Schittkowski's very successful Sequential Quadratic Programming NLPQLP software was obtained and used to experiment and analyze typical optimization problems of the type encountered in various rotorcraft wind tunnel and flight tests. Emphasis was directed toward obtaining efficiency, robustness and speed in computation.

Use of the NLPQLP↗

Preliminary study of the use of the STAR-100 computer for transonic flow calculations

An explicit method for solving the transonic small-disturbance potential equation is presented. This algorithm, which is suitable for the new vector-processor computers such as the CDC STAR-100, is compared to successive line over-relaxation (SLOR) on a simple test problem. The convergence rate of the explicit scheme is slower than that of SLOR, however, the efficiency of the explicit scheme on the STAR-100 computer is sufficient to overcome the slower convergence rate and allow an overall speedup compared to SLOR on the CYBER 175 computer.

Keller, J. D.↗

A feasibility study of a 3-D finite element solution scheme for aeroengine duct acoustics

The advantage from development of a 3-D model of aeroengine duct acoustics is the ability to analyze axial and circumferential liner segmentation simultaneously. The feasibility of a 3-D duct acoustics model was investigated using Galerkin or least squares element formulations combined with Gaussian elimination, successive over-relaxation, or conjugate gradient solution algorithms on conventional scalar computers and on a vector machine. A least squares element formulation combined with a conjugate gradient solver on a CDC Star vector computer initially appeared to have great promise, but severe difficulties were encountered with matrix ill-conditioning. These difficulties in conditioning rendered this technique impractical for realistic problems.

Abrahamson, A. L.↗

Dynamic scheduling of runway operations

Automated ATM/C decision making is discussed. Runway scheduling and flight plan generator algorithms are considered. Terminal area geometry, ATM/C schematics, vector controller display and simulation work are reported.

Pararas, J.↗

A highly optimized vectorized code for Monte Carlo simulations of SU(3) lattice gauge theories

New methods are introduced for improving the performance of the vectorized Monte Carlo SU(3) lattice gauge theory algorithm using the CDC CYBER 205. Structure, algorithm and programming considerations are discussed. The performance achieved for a 16(4) lattice on a 2-pipe system may be phrased in terms of the link update time or overall MFLOPS rates. For 32-bit arithmetic, it is 36.3 microsecond/link for 8 hits per iteration (40.9 microsecond for 10 hits) or 101.5 MFLOPS.

Barkai, D.↗

Robot arm dynamic model reduction for control

General methods are described by which the mathematical complexities of explicit and exact state equations of robot arms can be reduced to a simplified and compact state equation representation without introducing significant errors into the robot arm dynamic model. The model reduction methods are based on homogeneous coordinates and on the Langrangian algorithm for robot arm dynamics, and utilize matrix, vector and numeric analysis techniques. The derivation of differential vector representation of centripetal and Coriolis forces which has not yet been established in the literature is presented.

Bejczy, A. K.↗

An unsymmetric Lanczos algorithm for damped structural dynamics systems

A one-sided, unsymmetric block Lanczos algorithm is proposed for the model reduction of structural dynamics systems with unsymmetric damping and/or stiffness matrices. The algorithm is a three-term iteration scheme, which transforms the system matrix into an almost skew-symmetric, block-tridiagonal form. The Lanczos reduced-order model is guaranteed to be stable if the full-order system is stable. For unstable systems, a shifting method is available. Also, the algorithm offers flexibility in the choice of starting vectors and thus can yield more accurate reduced-order models. A linear system example and a plane truss structure example are used to show the efficacy of the proposed method.

Su, Tzu-Jeng↗

Determination of coronal magnetic fields from vector magnetograms

The determination of coronal magnetic fields from vector magnetograms, including the development and application of algorithms to determine force-free coronal fields above selected observations of active regions is studied. Two additional active regions were selected and analyzed. The restriction of periodicity in the 3-D code which is used to determine the coronal field was removed giving the new code variable mesh spacing and is thus able to provide a more realistic description of coronal fields. The NOAA active region AR5747 of 20 Oct. 1989 was studied. A brief account of progress during the research performed is reported.

Mikic, Zoran↗

Feed-forward volume rendering algorithm for moderately parallel MIMD machines

Algorithms for direct volume rendering on parallel and vector processors are investigated. Volumes are transformed efficiently on parallel processors by dividing the data into slices and beams of voxels. Equal sized sets of slices along one axis are distributed to processors. Parallelism is achieved at two levels. Because each slice can be transformed independently of others, processors transform their assigned slices with no communication, thus providing maximum possible parallelism at the first level. Within each slice, consecutive beams are incrementally transformed using coherency in the transformation computation. Also, coherency across slices can be exploited to further enhance performance. This coherency yields the second level of parallelism through the use of the vector processing or pipelining. Other ongoing efforts include investigations into image reconstruction techniques, load balancing strategies, and improving performance.

Yagel, Roni↗

Determination of coronal magnetic fields from vector magnetograms

This report covers technical progress during the second year of the contract entitled 'Determination of Coronal Magnetic Fields from Vector Magnetograms,' NASW-4728, between NASA and Science Applications International Corporation, and covers the period January 1, 1993 to December 31, 1993. Under this contract SAIC has conducted research into the determination of coronal magnetic fields from vector magnetograms, including the development and application of algorithms to determine force-free coronal fields above selected observations of active regions. The contract began on June 30, 1992 and has a completion date of December 31, 1994. This contract is a continuation of work started in a previous contract, NASW-4571, which covered the period November 15, 1990 to December 14, 1991. During this second year we have concentrated on studying additional active regions and in using the estimated coronal magnetic fields to compare to coronal features inferred from observations.

Mikic, Zoran↗

Image analysis by geostatistical and neural-network methods applications in glaciology

The applicability of neural network techniques, in the classification of ice surfaces and crevasse patterns, was analyzed. The observations of the Bering Glacier (Alaska) obtained from a surface survey and from the global positioning system (GPS) were used. A geographical information system was applied to test the usefulness of standard approaches. The information in the image needed to be reduced prior to the classification. The reduction was performed with a fast variogram algorithm sampling in three oblique directions. The resultant vectors provided the input for the neural network.

Herzfeld, Ute Christina↗

Numerical Simulations of Self-Focused Pulses Using the Nonlinear Maxwell Equations

This paper will present results in computational nonlinear optics. An algorithm will be described that solves the full vector nonlinear Maxwell's equations exactly without the approximations that are currently made. Present methods solve a reduced scalar wave equation, namely the nonlinear Schrodinger equation, and neglect the optical carrier. Also, results will be shown of calculations of 2-D electromagnetic nonlinear waves computed by directly integrating in time the nonlinear vector Maxwell's equations. The results will include simulations of 'light bullet' like pulses. Here diffraction and dispersion will be counteracted by nonlinear effects. The time integration efficiently implements linear and nonlinear convolutions for the electric polarization, and can take into account such quantum effects as Kerr and Raman interactions. The present approach is robust and should permit modeling 2-D and 3-D optical soliton propagation, scattering, and switching directly from the full-vector Maxwell's equations. Abstract of a proposed paper for presentation at the meeting NONLINEAR OPTICS: Materials, Fundamentals, and Applications, Hyatt Regency Waikaloa, Waikaloa, Hawaii, July 24-29, 1994, Cosponsored by IEEE/Lasers and Electro-Optics Society and Optical Society of America

Goorjian, Peter M.↗

A Radiation Chemistry Code Based on the Greens Functions of the Diffusion Equation

Ionizing radiation produces several radiolytic species such as.OH, e-aq, and H. when interacting with biological matter. Following their creation, radiolytic species diffuse and chemically react with biological molecules such as DNA. Despite years of research, many questions on the DNA damage by ionizing radiation remains, notably on the indirect effect, i.e. the damage resulting from the reactions of the radiolytic species with DNA. To simulate DNA damage by ionizing radiation, we are developing a step-by-step radiation chemistry code that is based on the Green's functions of the diffusion equation (GFDE), which is able to follow the trajectories of all particles and their reactions with time. In the recent years, simulations based on the GFDE have been used extensively in biochemistry, notably to simulate biochemical networks in time and space and are often used as the "gold standard" to validate diffusion-reaction theories. The exact GFDE for partially diffusion-controlled reactions is difficult to use because of its complex form. Therefore, the radial Green's function, which is much simpler, is often used. Hence, much effort has been devoted to the sampling of the radial Green's functions, for which we have developed a sampling algorithm This algorithm only yields the inter-particle distance vector length after a time step; the sampling of the deviation angle of the inter-particle vector is not taken into consideration. In this work, we show that the radial distribution is predicted by the exact radial Green's function. We also use a technique developed by Clifford et al. to generate the inter-particle vector deviation angles, knowing the inter-particle vector length before and after a time step. The results are compared with those predicted by the exact GFDE and by the analytical angular functions for free diffusion. This first step in the creation of the radiation chemistry code should help the understanding of the contribution of the indirect effect in the formation of DNA damage and double-strand breaks.

Plante, Ianik↗

Space-Based Remote Sensing of Atmospheric Aerosols: The Multi-Angle Spectro-Polarimetric Frontier

The review of optical instrumentation, forward modeling, and inverse problem solution for the polarimetric aerosol remote sensing from space is presented. The special emphasis is given to the description of current airborne and satellite imaging polarimeters and also to modern satellite aerosol retrieval algorithms based on the measurements of the Stokes vector of reflected solar light as detected on a satellite. Various underlying surface reflectance models are discussed and evaluated.

polarimetry↗

Efficacy of Code Optimization on Cache-based Processors

The current common wisdom in the U.S. is that the powerful, cost-effective supercomputers of tomorrow will be based on commodity (RISC) micro-processors with cache memories. Already, most distributed systems in the world use such hardware as building blocks. This shift away from vector supercomputers and towards cache-based systems has brought about a change in programming paradigm, even when ignoring issues of parallelism. Vector machines require inner-loop independence and regular, non-pathological memory strides (usually this means: non-power-of-two strides) to allow efficient vectorization of array operations. Cache-based systems require spatial and temporal locality of data, so that data once read from main memory and stored in high-speed cache memory is used optimally before being written back to main memory. This means that the most cache-friendly array operations are those that feature zero or unit stride, so that each unit of data read from main memory (a cache line) contains information for the next iteration in the loop. Moreover, loops ought to be 'fat', meaning that as many operations as possible are performed on cache data-provided instruction caches do not overflow and enough registers are available. If unit stride is not possible, for example because of some data dependency, then care must be taken to avoid pathological strides, just ads on vector computers. For cache-based systems the issues are more complex, due to the effects of associativity and of non-unit block (cache line) size. But there is more to the story. Most modern micro-processors are superscalar, which means that they can issue several (arithmetic) instructions per clock cycle, provided that there are enough independent instructions in the loop body. This is another argument for providing fat loop bodies. With these restrictions, it appears fairly straightforward to produce code that will run efficiently on any cache-based system. It can be argued that although some of the important computational algorithms employed at NASA Ames require different programming styles on vector machines and cache-based machines, respectively, neither architecture class appeared to be favored by particular algorithms in principle. Practice tells us that the situation is more complicated. This report presents observations and some analysis of performance tuning for cache-based systems. We point out several counterintuitive results that serve as a cautionary reminder that memory accesses are not the only factors that determine performance, and that within the class of cache-based systems, significant differences exist.

VanderWijngaart, Rob F.↗

FFTs in external or hierarchical memory

A description is given of advanced techniques for computing an ordered FFT on a computer with external or hierarchical memory. These algorithms (1) require as few as two passes through the external data set, (2) use strictly unit stride, long vector transfers between main memory and external storage, (3) require only a modest amount of scratch space in main memory, and (4) are well suited for vector and parallel computation. Performance figures are included for implementations of some of these algorithms on Cray supercomputers. Of interest is the fact that a main memory version outperforms the current Cray library FFT routines on the Cray-2, the Cray X-MP, and the Cray Y-MP systems. Using all eight processors on the Cray Y-MP, this main memory routine runs at nearly 2 Gflops.

Bailey, David H.↗

Multiple Uplinks Per Antenna (MUPA) Signal Acquisition Schemes

The Deep Space Network (DSN) currently makes use of the technique of Multiple Spacecraft per Antenna (MSPA) where a single antenna is used to track multiple spacecraft downlinks within its beam, such as in the case of multiple spacecraft orbiting Mars at 8.4 GHz (X-band). It is desired to extend this technique to the uplink where a single station is used to send a signal to multiple spacecraft in order to make more efficient use of ground resources. This would be applicable to numerous smallsat constellations being considered for future missions or to future spacecraft at Venus, Mars, or more distant destinations that are all within the half-power beamwidth of a single 34-m diameter antenna. In one scheme, each spacecraft’s command sequences would be time multiplexed onto a single uplink frequency. Each spacecraft would lock onto the uplink signal and would accept only commands intended for it via special identifier codes. Each spacecraft would also emit a downlink signal to the ground that is coherent with the uplink signal but would have its own allocated frequency channel and identifier information. A couple of key challenges associated with using this technique need to be addressed. Because of the single uplink frequency, coherent turnaround for two-way Doppler and ranging would not conform to established ratios, thus the radios employed by the spacecraft would need to be capable of variable turnaround ratios. In addition, because of the different orbits or spacecraft trajectories, the relative Doppler shifts and rates can be large with respect to the common uplink signal whose frequency would lie at the centroid of the frequencies of the expected received signals of the constellation. This would be problematic with standard analog spacecraft radios whose acquisition bandwidths are relatively small (~1.7 kHz) relative to the large frequency offsets (~100 kHz) expected using the single frequency uplink technique. With the advent of software defined radios (SDRs), signal frequency search algorithms can be utilized within the flight software and/or programmable hardware (e.g., FPGAs) that can easily acquire and track signals with large frequency offsets and varying dynamics. Such techniques could include FFT search algorithms, step-and-sweep search algorithms, or onboard frequency steering making use of trajectory vectors uplinked to each member spacecraft. Other challenges include mitigation of potential interference between received signals. We have identified several software defined radios that are in different stages of development and whose key parameters have been tabulated. We have examined each radio’s capabilities with respect to acquiring and tracking signals with large frequency offsets. Such analyses made use of previous studies supplemented with specially designed tests using both simulation tools and/or existing testbeds. We have compared signal acquisition times computed from provided algorithms along with measured values derived from tests using existing hardware and simulation tools for the purpose of conducting tradeoff studies between the various radio designs and software/firmware programming approaches.

Abraham, Douglas S.↗