Search NASA⌕ Search

SEARCH · Search NASA

Results for “Vectorized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

A highly optimized vectorized code for Monte Carlo simulations of SU(3) lattice gauge theories

New methods are introduced for improving the performance of the vectorized Monte Carlo SU(3) lattice gauge theory algorithm using the CDC CYBER 205. Structure, algorithm and programming considerations are discussed. The performance achieved for a 16(4) lattice on a 2-pipe system may be phrased in terms of the link update time or overall MFLOPS rates. For 32-bit arithmetic, it is 36.3 microsecond/link for 8 hits per iteration (40.9 microsecond for 10 hits) or 101.5 MFLOPS.

Barkai, D.↗

Robot arm dynamic model reduction for control

General methods are described by which the mathematical complexities of explicit and exact state equations of robot arms can be reduced to a simplified and compact state equation representation without introducing significant errors into the robot arm dynamic model. The model reduction methods are based on homogeneous coordinates and on the Langrangian algorithm for robot arm dynamics, and utilize matrix, vector and numeric analysis techniques. The derivation of differential vector representation of centripetal and Coriolis forces which has not yet been established in the literature is presented.

Bejczy, A. K.↗

An unsymmetric Lanczos algorithm for damped structural dynamics systems

A one-sided, unsymmetric block Lanczos algorithm is proposed for the model reduction of structural dynamics systems with unsymmetric damping and/or stiffness matrices. The algorithm is a three-term iteration scheme, which transforms the system matrix into an almost skew-symmetric, block-tridiagonal form. The Lanczos reduced-order model is guaranteed to be stable if the full-order system is stable. For unstable systems, a shifting method is available. Also, the algorithm offers flexibility in the choice of starting vectors and thus can yield more accurate reduced-order models. A linear system example and a plane truss structure example are used to show the efficacy of the proposed method.

Su, Tzu-Jeng↗

Determination of coronal magnetic fields from vector magnetograms

The determination of coronal magnetic fields from vector magnetograms, including the development and application of algorithms to determine force-free coronal fields above selected observations of active regions is studied. Two additional active regions were selected and analyzed. The restriction of periodicity in the 3-D code which is used to determine the coronal field was removed giving the new code variable mesh spacing and is thus able to provide a more realistic description of coronal fields. The NOAA active region AR5747 of 20 Oct. 1989 was studied. A brief account of progress during the research performed is reported.

Mikic, Zoran↗

Feed-forward volume rendering algorithm for moderately parallel MIMD machines

Algorithms for direct volume rendering on parallel and vector processors are investigated. Volumes are transformed efficiently on parallel processors by dividing the data into slices and beams of voxels. Equal sized sets of slices along one axis are distributed to processors. Parallelism is achieved at two levels. Because each slice can be transformed independently of others, processors transform their assigned slices with no communication, thus providing maximum possible parallelism at the first level. Within each slice, consecutive beams are incrementally transformed using coherency in the transformation computation. Also, coherency across slices can be exploited to further enhance performance. This coherency yields the second level of parallelism through the use of the vector processing or pipelining. Other ongoing efforts include investigations into image reconstruction techniques, load balancing strategies, and improving performance.

Yagel, Roni↗

Determination of coronal magnetic fields from vector magnetograms

This report covers technical progress during the second year of the contract entitled 'Determination of Coronal Magnetic Fields from Vector Magnetograms,' NASW-4728, between NASA and Science Applications International Corporation, and covers the period January 1, 1993 to December 31, 1993. Under this contract SAIC has conducted research into the determination of coronal magnetic fields from vector magnetograms, including the development and application of algorithms to determine force-free coronal fields above selected observations of active regions. The contract began on June 30, 1992 and has a completion date of December 31, 1994. This contract is a continuation of work started in a previous contract, NASW-4571, which covered the period November 15, 1990 to December 14, 1991. During this second year we have concentrated on studying additional active regions and in using the estimated coronal magnetic fields to compare to coronal features inferred from observations.

Mikic, Zoran↗

Image analysis by geostatistical and neural-network methods applications in glaciology

The applicability of neural network techniques, in the classification of ice surfaces and crevasse patterns, was analyzed. The observations of the Bering Glacier (Alaska) obtained from a surface survey and from the global positioning system (GPS) were used. A geographical information system was applied to test the usefulness of standard approaches. The information in the image needed to be reduced prior to the classification. The reduction was performed with a fast variogram algorithm sampling in three oblique directions. The resultant vectors provided the input for the neural network.

Herzfeld, Ute Christina↗

Numerical Simulations of Self-Focused Pulses Using the Nonlinear Maxwell Equations

This paper will present results in computational nonlinear optics. An algorithm will be described that solves the full vector nonlinear Maxwell's equations exactly without the approximations that are currently made. Present methods solve a reduced scalar wave equation, namely the nonlinear Schrodinger equation, and neglect the optical carrier. Also, results will be shown of calculations of 2-D electromagnetic nonlinear waves computed by directly integrating in time the nonlinear vector Maxwell's equations. The results will include simulations of 'light bullet' like pulses. Here diffraction and dispersion will be counteracted by nonlinear effects. The time integration efficiently implements linear and nonlinear convolutions for the electric polarization, and can take into account such quantum effects as Kerr and Raman interactions. The present approach is robust and should permit modeling 2-D and 3-D optical soliton propagation, scattering, and switching directly from the full-vector Maxwell's equations. Abstract of a proposed paper for presentation at the meeting NONLINEAR OPTICS: Materials, Fundamentals, and Applications, Hyatt Regency Waikaloa, Waikaloa, Hawaii, July 24-29, 1994, Cosponsored by IEEE/Lasers and Electro-Optics Society and Optical Society of America

Goorjian, Peter M.↗

A Radiation Chemistry Code Based on the Greens Functions of the Diffusion Equation

Ionizing radiation produces several radiolytic species such as.OH, e-aq, and H. when interacting with biological matter. Following their creation, radiolytic species diffuse and chemically react with biological molecules such as DNA. Despite years of research, many questions on the DNA damage by ionizing radiation remains, notably on the indirect effect, i.e. the damage resulting from the reactions of the radiolytic species with DNA. To simulate DNA damage by ionizing radiation, we are developing a step-by-step radiation chemistry code that is based on the Green's functions of the diffusion equation (GFDE), which is able to follow the trajectories of all particles and their reactions with time. In the recent years, simulations based on the GFDE have been used extensively in biochemistry, notably to simulate biochemical networks in time and space and are often used as the "gold standard" to validate diffusion-reaction theories. The exact GFDE for partially diffusion-controlled reactions is difficult to use because of its complex form. Therefore, the radial Green's function, which is much simpler, is often used. Hence, much effort has been devoted to the sampling of the radial Green's functions, for which we have developed a sampling algorithm This algorithm only yields the inter-particle distance vector length after a time step; the sampling of the deviation angle of the inter-particle vector is not taken into consideration. In this work, we show that the radial distribution is predicted by the exact radial Green's function. We also use a technique developed by Clifford et al. to generate the inter-particle vector deviation angles, knowing the inter-particle vector length before and after a time step. The results are compared with those predicted by the exact GFDE and by the analytical angular functions for free diffusion. This first step in the creation of the radiation chemistry code should help the understanding of the contribution of the indirect effect in the formation of DNA damage and double-strand breaks.

Plante, Ianik↗

Space-Based Remote Sensing of Atmospheric Aerosols: The Multi-Angle Spectro-Polarimetric Frontier

The review of optical instrumentation, forward modeling, and inverse problem solution for the polarimetric aerosol remote sensing from space is presented. The special emphasis is given to the description of current airborne and satellite imaging polarimeters and also to modern satellite aerosol retrieval algorithms based on the measurements of the Stokes vector of reflected solar light as detected on a satellite. Various underlying surface reflectance models are discussed and evaluated.

polarimetry↗

Efficacy of Code Optimization on Cache-based Processors

The current common wisdom in the U.S. is that the powerful, cost-effective supercomputers of tomorrow will be based on commodity (RISC) micro-processors with cache memories. Already, most distributed systems in the world use such hardware as building blocks. This shift away from vector supercomputers and towards cache-based systems has brought about a change in programming paradigm, even when ignoring issues of parallelism. Vector machines require inner-loop independence and regular, non-pathological memory strides (usually this means: non-power-of-two strides) to allow efficient vectorization of array operations. Cache-based systems require spatial and temporal locality of data, so that data once read from main memory and stored in high-speed cache memory is used optimally before being written back to main memory. This means that the most cache-friendly array operations are those that feature zero or unit stride, so that each unit of data read from main memory (a cache line) contains information for the next iteration in the loop. Moreover, loops ought to be 'fat', meaning that as many operations as possible are performed on cache data-provided instruction caches do not overflow and enough registers are available. If unit stride is not possible, for example because of some data dependency, then care must be taken to avoid pathological strides, just ads on vector computers. For cache-based systems the issues are more complex, due to the effects of associativity and of non-unit block (cache line) size. But there is more to the story. Most modern micro-processors are superscalar, which means that they can issue several (arithmetic) instructions per clock cycle, provided that there are enough independent instructions in the loop body. This is another argument for providing fat loop bodies. With these restrictions, it appears fairly straightforward to produce code that will run efficiently on any cache-based system. It can be argued that although some of the important computational algorithms employed at NASA Ames require different programming styles on vector machines and cache-based machines, respectively, neither architecture class appeared to be favored by particular algorithms in principle. Practice tells us that the situation is more complicated. This report presents observations and some analysis of performance tuning for cache-based systems. We point out several counterintuitive results that serve as a cautionary reminder that memory accesses are not the only factors that determine performance, and that within the class of cache-based systems, significant differences exist.

VanderWijngaart, Rob F.↗

FFTs in external or hierarchical memory

A description is given of advanced techniques for computing an ordered FFT on a computer with external or hierarchical memory. These algorithms (1) require as few as two passes through the external data set, (2) use strictly unit stride, long vector transfers between main memory and external storage, (3) require only a modest amount of scratch space in main memory, and (4) are well suited for vector and parallel computation. Performance figures are included for implementations of some of these algorithms on Cray supercomputers. Of interest is the fact that a main memory version outperforms the current Cray library FFT routines on the Cray-2, the Cray X-MP, and the Cray Y-MP systems. Using all eight processors on the Cray Y-MP, this main memory routine runs at nearly 2 Gflops.

Bailey, David H.↗

Multiple Uplinks Per Antenna (MUPA) Signal Acquisition Schemes

The Deep Space Network (DSN) currently makes use of the technique of Multiple Spacecraft per Antenna (MSPA) where a single antenna is used to track multiple spacecraft downlinks within its beam, such as in the case of multiple spacecraft orbiting Mars at 8.4 GHz (X-band). It is desired to extend this technique to the uplink where a single station is used to send a signal to multiple spacecraft in order to make more efficient use of ground resources. This would be applicable to numerous smallsat constellations being considered for future missions or to future spacecraft at Venus, Mars, or more distant destinations that are all within the half-power beamwidth of a single 34-m diameter antenna. In one scheme, each spacecraft’s command sequences would be time multiplexed onto a single uplink frequency. Each spacecraft would lock onto the uplink signal and would accept only commands intended for it via special identifier codes. Each spacecraft would also emit a downlink signal to the ground that is coherent with the uplink signal but would have its own allocated frequency channel and identifier information. A couple of key challenges associated with using this technique need to be addressed. Because of the single uplink frequency, coherent turnaround for two-way Doppler and ranging would not conform to established ratios, thus the radios employed by the spacecraft would need to be capable of variable turnaround ratios. In addition, because of the different orbits or spacecraft trajectories, the relative Doppler shifts and rates can be large with respect to the common uplink signal whose frequency would lie at the centroid of the frequencies of the expected received signals of the constellation. This would be problematic with standard analog spacecraft radios whose acquisition bandwidths are relatively small (~1.7 kHz) relative to the large frequency offsets (~100 kHz) expected using the single frequency uplink technique. With the advent of software defined radios (SDRs), signal frequency search algorithms can be utilized within the flight software and/or programmable hardware (e.g., FPGAs) that can easily acquire and track signals with large frequency offsets and varying dynamics. Such techniques could include FFT search algorithms, step-and-sweep search algorithms, or onboard frequency steering making use of trajectory vectors uplinked to each member spacecraft. Other challenges include mitigation of potential interference between received signals. We have identified several software defined radios that are in different stages of development and whose key parameters have been tabulated. We have examined each radio’s capabilities with respect to acquiring and tracking signals with large frequency offsets. Such analyses made use of previous studies supplemented with specially designed tests using both simulation tools and/or existing testbeds. We have compared signal acquisition times computed from provided algorithms along with measured values derived from tests using existing hardware and simulation tools for the purpose of conducting tradeoff studies between the various radio designs and software/firmware programming approaches.

Abraham, Douglas S.↗

Robust integration schemes for generalized viscoplasticity with internal-state variables. Part 2: Algorithmic developments and implementation

This two-part report is concerned with the development of a general framework for the implicit time-stepping integrators for the flow and evolution equations in generalized viscoplastic models. The primary goal is to present a complete theoretical formulation, and to address in detail the algorithmic and numerical analysis aspects involved in its finite element implementation, as well as to critically assess the numerical performance of the developed schemes in a comprehensive set of test cases. On the theoretical side, the general framework is developed on the basis of the unconditionally-stable, backward-Euler difference scheme as a starting point. Its mathematical structure is of sufficient generality to allow a unified treatment of different classes of viscoplastic models with internal variables. In particular, two specific models of this type, which are representative of the present start-of-art in metal viscoplasticity, are considered in applications reported here; i.e., fully associative (GVIPS) and non-associative (NAV) models. The matrix forms developed for both these models are directly applicable for both initially isotropic and anisotropic materials, in general (three-dimensional) situations as well as subspace applications (i.e., plane stress/strain, axisymmetric, generalized plane stress in shells). On the computational side, issues related to efficiency and robustness are emphasized in developing the (local) interative algorithm. In particular, closed-form expressions for residual vectors and (consistent) material tangent stiffness arrays are given explicitly for both GVIPS and NAV models, with their maximum sizes 'optimized' to depend only on the number of independent stress components (but independent of the number of viscoplastic internal state parameters). Significant robustness of the local iterative solution is provided by complementing the basic Newton-Raphson scheme with a line-search strategy for convergence. In the present second part of the report, we focus on the specific details of the numerical schemes, and associated computer algorithms, for the finite-element implementation of GVIPS and NAV models.

Li, Wei↗

Assessment of Polarization Effect on Efficiency of Levenberg-Marquardt Algorithm in Case of Thin Atmosphere Over Black Surface

The Levenberg-Marquardt algorithm [1, 2] provides a numerical iterative solution to the problem of minimization of a function over a space of its parameters. In our work, the Levenberg-Marquardt algorithm retrieves optical parameters of a thin (single scattering) plane parallel atmosphere irradiated by collimated infinitely wide monochromatic beam of light. Black ground surface is assumed. Computational accuracy, sensitivity to the initial guess and the presence of noise in the signal, and other properties of the algorithm are investigated in scalar (using intensity only) and vector (including polarization) modes. We consider an atmosphere that contains a mixture of coarse and fine fractions. Following [3], the fractions are simulated using Henyey-Greenstein model. Though not realistic, this assumption is very convenient for tests [4, p.354]. In our case it yields analytical evaluation of Jacobian matrix. Assuming the MISR geometry of observation [5] as an example, the average scattering cosines and the ratio of coarse and fine fractions, the atmosphere optical depth, and the single scattering albedo, are the five parameters to be determined numerically. In our implementation of the algorithm, the system of five linear equations is solved using the fast Cramer s rule [6]. A simple subroutine developed by the authors, makes the algorithm independent from external libraries. All Fortran 90/95 codes discussed in the presentation will be available immediately after the meeting from sergey.v.korkin@nasa.gov by request.

Korkin, S.↗

A Fast Vector Radiative Transfer Model for the Atmosphere-Ocean Coupled System

To infer atmospheric and oceanic constituent properties from polarimetric observations, an efficient and accurate retrieval algorithm is desirable. In-line radiative transfer calculations are indispensable if a large state vector, including both atmospheric profiles and surface properties, is used to improve retrieval accuracy. However, in-line radiative transfer calculations are usually not computationally efficient for remote sensing applications. Therefore, there is a pressing need to develop an accurate and fast vector radiative transfer model (RTM) to fully utilize satellite polarimetric observations. This paper reports on a fast vector RTM, referred to as TAMU-VRTM, in support of polarimetric remote sensing, which is capable of simulating the Stokes vector values observed at the top of the atmosphere and at the surface by fully considering absorption, scattering, and emission in the atmosphere and ocean. Gaseous absorption is parameterized with respect to gas concentration, temperature, and pressure, by using a regression method applicable to an inhomogeneous atmospheric path. An efficient two-component approach combining the small angle approximation and the adding-doubling method is utilized to solve the vector radiative transfer equation (RTE). The thermal emission component of the RTE solution is obtained by an efficient doubling process. The air-sea interface is treated as a wind-ruffled rough surface in the model to mimic a realistic ocean surface. Several oceanic optical property models are introduced to model ocean inherent optical properties. To demonstrate the applicability of the TAMUVRTM, simulations are compared with satellite observations, and results from other vector radiative transfer methods including benchmarks.

radiative transfer in coupled atmosphere-ocean sys↗

MeV Gamma Ray Detection Algorithms for Stacked Silicon Detectors

By making use of the signature of a gamma ray event as it appears in N = 5 to 20 lithium-drifted silicon detectors and applying smart selection algorithms, gamma rays in the energy range of 1 to 8 MeV can be detected with good efficiency and selectivity. Examples of the types of algorithms used for different energy regions include the simple sum mode, the sum-coincidence mode used in segmented detectors, unique variations on sum-coincidence for an N-dimensional vector event, and a new and extremely useful mode for double escape peak spectroscopy at pair-production energies. The latter algorithm yields a spectrum similar to that of the pair spectrometer, but without the need of the dual external segments for double escape coincidence, and without the large loss in efficiency of double escape events. Background events due to Compton scattering are largely suppressed. Monte Carlo calculations were used to model the gamma ray interactions in the silicon, in order to enable testing of a wide array of different algorithms on the event N-vectors for a large-N stack.

McMurray, Robert E. Jr.↗

Applications of the Dynamic N-Dimensional K-Vector

The n-dimensional k-vector (NDKV) is an appealing alternative to binary tress for resolving complex queries in large relational databases. The method has excelled in several applications involving static databases. The present paper extends the theory supporting the NDKV to handle dynamic databases, where the data is updated frequently. This includes deleting records, adding new entries, or editing existing elements. The merit of this new version of the NDKV, the dynamic n-dimensional k-vector (DNDKV), is that it is no longer necessary to recompute the entire k-vector (the main structure that indexes the data) every time a record changes. The algorithm updates the four constituents of the standard NDKV on the fly: the database, sorted database, index, and k-vector tables. As a result, the DNDKV becomes comparable in terms of capabilities and flexibility to stateof-the-art storage engines relying on structured query languages (SQL). The performance of the DNDKV is assessed by running typical read/write operations on a database that contains millions of pre-computed missions to celestial bodies. This database requires frequent updates whenever an orbit solution is refined or new bodies are discovered. The DNDKV is faster than rebuilding the k-vector tables completely, provided that the number of elements being added or removed is not excessively large. Direct runtime comparisons with MySQL suggest that the DNDKV is several times faster for reading but might be slower for writing and updating the database. One limit of the technique is the elements being added must be within the range of the current k-vector tables. If this is not the case, the technique cannot be used and the k-vector tables must be rebuilt from scratch.

Mortari, Daniele↗