Search NASA⌕ Search

SEARCH · Search NASA

Results for “Vectorized algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

Distribution of mica polytypes among space groups.

All the possible space groups for mica polytypes are deduced by making use of the characteristics of the mica unit layer and stacking mode. The algebraic properties of the vector-stacking symbol of Ross et al. (1966) are examined, and a simple algorithm for deducing the space group from this symbol is presented. A method considered for enumerating all possible stacking sequences of mica polytypes makes use of a computer.

Takeda, H.↗

Three-dimensional computational aerodynamics in the 1980's

The future requirements for constructing codes that can be used to compute three-dimensional flows about aerodynamic shapes should be assessed in light of the constraints imposed by future computer architectures and the reality of usable algorithms that can provide practical three-dimensional simulations. On the hardware side, vector processing is inevitable in order to meet the CPU speeds required. To cope with three-dimensional geometries, massive data bases with fetch/store conflicts and transposition problems are inevitable. On the software side, codes must be prepared that: (1) can be adapted to complex geometries, (2) can (at the very least) predict the location of laminar and turbulent boundary layer separation, and (3) will converge rapidly to sufficiently accurate solutions.

Lomax, H.↗

A vectorization of the Jameson-Caughey NYU transonic swept-wing computer program FLO-22-V1 for the STAR-100 computer

The computer program FLO-22 for analyzing inviscid transonic flow past 3-D swept-wing configurations was modified to use vector operations and run on the STAR-100 computer. The vectorized version described herein was called FLO-22-V1. Vector operations were incorporated into Successive Line Over-Relaxation in the transformed horizontal direction. Vector relational operations and control vectors were used to implement upwind differencing at supersonic points. A high speed of computation and extended grid domain were characteristics of FLO-22-V1. The new program was not the optimal vectorization of Successive Line Over-Relaxation applied to transonic flow; however, it proved that vector operations can readily be implemented to increase the computation rate of the algorithm.

Smith, R. E.↗

A Taylor-Galerkin finite element algorithm for transient nonlinear thermal-structural analysis

A Taylor-Galerkin finite element method for solving large, nonlinear thermal-structural problems is presented. The algorithm is formulated for coupled transient and uncoupled quasistatic thermal-structural problems. Vectorizing strategies ensure computational efficiency. Two applications demonstrate the validity of the approach for analyzing transient and quasistatic thermal-structural problems.

Thornton, E. A.↗

A Taylor-Galerkin finite element algorithm for transient nonlinear thermal-structural analysis

A Taylor-Galerkin finite element method for solving large, nonlinear thermal-structural problems is presented. The algorithm is formulated for coupled transient and uncoupled quasistatic thermal-structural problems. Vectorizing strategies ensure computational efficiency. Two applications demonstrate the validity of the approach for analyzing transient and quasistatic thermal-structural problems.

Thornton, E. A.↗

Comparison between the PISO algorithm and preconditioning methods for compressible flow

Two widely used family of algorithms, pressure-based and density-based methods, have been developed for computational fluid dynamics (CFD) problems over the years. Pressure-based methods (such as SIMPLE and PISO) use a Poisson-like equation for updating pressure instead of the continuity equation, while density-based methods use the continuity equation to update density (an equation of state is used to provide density in pressure based schemes and pressure in density based schemes). Pressure-based methods were developed originally for incompressible flows at low Reynolds numbers and were then extended to high Reynolds numbers and compressible applications. On the other hand, density based methods were originally developed for transonic flows and have been extended down to low Mach numbers through the use of preconditioning techniques. We compare these two very different approaches to solving the Navier-Stokes equations in order to gain an understanding of their similarities and differences. Specifically, we consider the PISO scheme as a representative pressure-based method and contrast it with a recently developed preconditioning scheme. We also compare the relative performance of the PISO algorithm with a Euler implicit algorithm that is employed to solve the preconditioned equations by means of a vector stability analysis.

Merkle, Charles L.↗

Numerical Procedures for Inlet/Diffuser/Nozzle Flows

Two primitive variable, pressure based, flux-split, RNS/NS solution procedures for viscous flows are presented. Both methods are uniformly valid across the full Mach number range, Le., from the incompressible limit to high supersonic speeds. The first method is an 'optimized' version of a previously developed global pressure relaxation RNS procedure. Considerable reduction in the number of relatively expensive matrix inversion, and thereby in the computational time, has been achieved with this procedure. CPU times are reduced by a factor of 15 for predominantly elliptic flows (incompressible and low subsonic). The second method is a time-marching, 'linearized' convection RNS/NS procedure. The key to the efficiency of this procedure is the reduction to a single LU inversion at the inflow cross-plane. The remainder of the algorithm simply requires back-substitution with this LU and the corresponding residual vector at any cross-plane location. This method is not time-consistent, but has a convective-type CFL stability limitation. Both formulations are robust and provide accurate solutions for a variety of internal viscous flows to be provided herein.

Rubin, Stanley G.↗

Statistical Study of the Properties of Magnetosheath Lion Roars

Lion roars are narrowband whistler wave emissions that have been observed in several environments, such as planetary magnetosheaths, the Earth's magnetosphere, the solar wind, downstream of interplanetary shocks, and the cusp region. We present measurements of more than 30,000 such emissions observed by the Magnetospheric Multiscale spacecraft with high‐cadence (8,192 samples/s) search coil magnetometer data. A semiautomatic algorithm was used to identify the emissions, and an adaptive interval algorithm in conjunction with minimum variance analysis was used to determine their wave vector. The properties of the waves are determined in both the spacecraft and plasma rest frame. The mean wave normal angle, with respect to the background magnetic field (B(sub 0)), plasma bulk flow velocity (V(sub b)), and the coplanarity plane (V(sub b) × B(sub 0)) are 23°, 56°, and 0°, respectively. The average peak frequencies were ∼31% of the electron gyrofrequency (ω(sub ce)) observed in the spacecraft frame and ∼18% of ω(sub ce) in the plasma rest frame. In the spacecraft frame, ∼99% of the emissions had a frequency <ω(sub ce), while 98% had a peak frequency <0.72 ω(sub ce) in the plasma rest frame. None of the waves had frequencies lower than the lower hybrid frequency, ω. From the probability density function of the electron plasma β(sub e), the ratio between the electron thermal and magnetic pressure, ∼99.6% of the waves were observed with β(sub e)<4 with a large narrow peak at 0.07 and two smaller, but wider, peaks at 1.26 and 2.28, while the average value was ∼1.25.

Magnetosheath emissions↗

First-Order Runtime Verification using BDDs

Runtime Verification (RV) expedites the analyses of execution traces for detecting system errors and for statistical and quality analysis. Having started modestly, with checking temporal properties that are based on propositional (yes/no) values, the current practice of RV often involves properties that are parametrized by the data observed in the input trace. The specifications are based on various formalisms, such as automata, temporal logics, rule systems, and stream processing. Checking execution traces that are data intensive against a specification that imposes strong dependencies between the data, poses a nontrivial challenges; in particular if runtime verification has to be performed online, while many events that carry data appear within small time proximities. Towards achieving this goal, it was recently suggested to represent relations over the observed data values, based on BDDs, where data elements are enumerated and then converted into bit vectors. This representation provided a very simple and natural extension of an RV algorithm from propositional to first-order LTL, but more importantly, was shown to contribute to the memory compactness and to the speed, as was demonstrated using a corresponding implementation. We extend here the capabilities of BDD-based RV with the ability to express timing constraints, where the monitored events include (integer) clock values. We show how to efficiently operate on BDDs that represent both relations on (enumerations of) values and time dependencies, as required by the addition of the time constraints. We demonstrate our algorithm with an efficient implementation and provide experimental results.

Peled, Doron↗

A Design Study of Onboard Navigation and Guidance During Aerocapture at Mars

The navigation and guidance of a high lift-to-drag ratio sample return vehicle during aerocapture at Mars are investigated. Emphasis is placed on integrated systems design, with guidance algorithm synthesis and analysis based on vehicle state and atmospheric density uncertainty estimates provided by the navigation system. The latter utilizes a Kalman filter for state vector estimation, with useful update information obtained through radar altimeter measurements and density altitude measurements based on IMU-measured drag acceleration. A three-phase guidance algorithm, featuring constant bank numeric predictor/corrector atmospheric capture and exit phases and an extended constant altitude cruise phase, is developed to provide controlled capture and depletion of orbital energy, orbital plane control, and exit apoapsis control. Integrated navigation and guidance systems performance are analyzed using a four degree-of-freedom computer simulation. The simulation environment includes an atmospheric density model with spatially correlated perturbations to provide realistic variations over the vehicle trajectory. Navigation filter initial conditions for the analysis are based on planetary approach optical navigation results. Results from a selection of test cases are presented to give insight into systems performance.

Fuhry, Douglas Paul↗

Skyline based terrain matching

Skyline-based terrain matching, a new method for locating the vantage point of stereo camera or laser range-finding measurements on a global map previously prepared by satellite or aerial mapping is described. The orientation of the vantage is assumed known, but its translational parameters are determined by the algorithm. Skylines, or occluding contours, can be extracted from the sensory measurements taken by an autonomous vehicle. They can also be modeled from the global map, given a vantage estimate from which to start. The two sets of skylines, represented in cylindrical coordinates about either the true or the estimated vantage, are employed as 'features' or reference objects common to both sources of information. The terrain matching problem is formulated in terms of finding a translation between the respective representations of the skylines, by approximating the two sets of skylines as identical features (curves) on the actual terrain. The search for this translation is based on selecting the longest of the minimum-distance vectors between corresponding curves from the two sets of skylines. In successive iterations of the algorithm, the approximation that the two sets of curves are identical becomes more accurate, and the vantage estimate continues to improve. The algorithm was implemented and evaluated on a simulated terrain. Illustrations and examples are included.

Page, Lance A.↗

Accelerating GNNs on GPU Sparse Tensor Cores through N:M Sparsity-Oriented Graph Reordering

Recent advancements in GPU hardware support have introduced the capability to leverage N:M sparse patterns for substantial performance gains. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to such sparse patterns. In this paper, we propose a novel graph reordering algorithm, the first of its kind, to reshape irregular graph data into the N:M structured sparse pattern at the tile level, allowing linear-algebra-based graph operations in GNNs to benefit from the N:M sparse hardware. The optimization is lossless, maintaining the accuracy of GNN. It can remove 98-100\% violations of the N:M sparse patterns at the vector level, and increase the proportion of conforming graphs in SuiteSparse collection from 5-9\% to 88.7-93.5\%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (2.3X -- 7.5X on average) and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average).

artificial intelligence, graph neural networks↗

Simplified Decoding of Convolutional Codes

Some complicated intermediate steps shortened or eliminated. Decoding of convolutional error-correcting digital codes simplified by new errortrellis syndrome technique. In new technique, syndrome vector not computed. Instead, advantage taken of newly-derived mathematical identities simplify decision tree, folding it back on itself into form called "error trellis." This trellis graph of all path solutions of syndrome equations. Each path through trellis corresponds to specific set of decisions as to received digits. Existing decoding algorithms combined with new mathematical identities reduce number of combinations of errors considered and enable computation of correction vector directly from data and check bits as received.

Truong, T. K.↗

Newton algorithm for fitting transfer functions to frequency response measurements

In this paper the problem of synthesizing transfer functions from frequency response measurements is considered. Given a complex vector representing the measured frequency response of a physical system, a transfer function of specified order is determined that minimizes the sum of the magnitude-squared of the frequency response errors. This nonlinear least squares minimization problem is solved by an iterative global descent algorithm of the Newton type that converges quadratically near the minimum. The unknown transfer function is expressed as a sum of second-order rational polynomials, a parameterization that facilitates a numerically robust computer implementation. The algorithm is developed for single-input, single-output, causal, stable transfer functions. Two numerical examples demonstrate the effectiveness of the algorithm.

Spanos, J. T.↗

A Local Macroscopic Conservative (LoMaC) Low Rank Tensor Method for the Vlasov Dynamics

Abstract In this paper, we propose a novel Local Macroscopic Conservative (LoMaC) low rank tensor method for simulating the Vlasov-Poisson (VP) system. The LoMaC property refers to the exact local conservation of macroscopic mass, momentum and energy at the discrete level. This is a follow-up work of our previous development of a conservative low rank tensor approach for Vlasov dynamics ( arXiv:2201.10397 ). In that work, we applied a low rank tensor method with a conservative singular value decomposition to the high dimensional VP system to mitigate the curse of dimensionality, while maintaining the local conservation of mass and momentum. However, energy conservation is not guaranteed, which is a critical property to avoid unphysical plasma self-heating or cooling. The new ingredient in the LoMaC low rank tensor algorithm is that we simultaneously evolve the macroscopic conservation laws of mass, momentum and energy using a flux-difference form with kinetic flux vector splitting; then the LoMaC property is realized by projecting the low rank kinetic solution onto a subspace that shares the same macroscopic observables by a conservative orthogonal projection. The algorithm is extended to the high dimensional problems by hierarchical Tuck decomposition of solution tensors and a corresponding conservative projection algorithm. Extensive numerical tests on the VP system are showcased for the algorithm’s efficacy.

Guo, Wei↗

The gust-front detection and wind-shift algorithms for the Terminal Doppler Weather Radar system

The Federal Aviation Administration's (FAA) Terminal Doppler Weather Radar (TDWR) system was primarily designed to address the operational needs of pilots in the avoidance of low-altitude wind shears upon takeoff and landing at airports. One of the primary methods of wind-shear detection for the TDWR system is the gust-front detection algorithm. The algorithm is designed to detect gust fronts that produce a wind-shear hazard and/or sustained wind shifts. It serves the hazard warning function by providing an estimate of the wind-speed gain for aircraft penetrating the gust front. The gust-front detection and wind-shift algorithms together serve a planning function by providing forecasted gust-front locations and estimates of the horizontal wind vector behind the front, respectively. This information is used by air traffic managers to determine arrival and departure runway configurations and aircraft movements to minimize the impact of wind shifts on airport capacity. This paper describes the gust-front detection and wind-shift algorithms to be fielded in the initial TDWR systems. Results of a quantitative performance evaluation using Doppler radar data collected during TDWR operational demonstrations at the Denver, Kansas City, and Orlando airports are presented. The algorithms were found to be operationally useful by the FAA airport controllers and supervisors.

Hermes, Laurie G.↗

State-Space System Realization with Input- and Output-Data Correlation

This paper introduces a general version of the information matrix consisting of the autocorrelation and cross-correlation matrices of the shifted input and output data. Based on the concept of data correlation, a new system realization algorithm is developed to create a model directly from input and output data. The algorithm starts by computing a special type of correlation matrix derived from the information matrix. The special correlation matrix provides information on the system-observability matrix and the state-vector correlation. A system model is then developed from the observability matrix in conjunction with other algebraic manipulations. This approach leads to several different algorithms for computing system matrices for use in representing the system model. The relationship of the new algorithms with other realization algorithms in the time and frequency domains is established with matrix factorization of the information matrix. Several examples are given to illustrate the validity and usefulness of these new algorithms.

Juang, Jer-Nan↗

Time scheduling of a mix of 4D equipped and unequipped aircraft

In planning for a future automated air traffic system, it is necessary to confront the transition situation in which some percentage of the traffic must be handled by conventional means. A safe, efficient transition system is needed since initially not all aircraft will be able to respond to a more automated system. The specific problem addressed was that of time scheduling a mix of 4D-equipped aircraft (aircraft that can accurately meet a controller specified time schedule at selected way points in the terminal area) when operating in conjunction with unequipped aircraft (aircraft that require air traffic handling by means of standard vectoring techniques). First, a relationship between time separation and system capacity was developed. The time separations were incorporated into a set of scheduling algorithms which contain the required elements of flexibility needed for terminal-area operation, such as delaying aircraft and changing time separations. The problem of reducing the size of time separations allotted for vectored aircraft by means of computer assists to the controller was also addressed.

Tobias, L.↗