Search NASA⌕ Search

SEARCH · Search NASA

Results for “Partitioned scheme”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Porting Gravitational Wave Signal Extraction to Parallel Virtual Machine (PVM)

Laser Interferometer Space Antenna (LISA) is a planned NASA-ESA mission to be launched around 2012. The Gravitational Wave detection is fundamentally the determination of frequency, source parameters, and waveform amplitude derived in a specific order from the interferometric time-series of the rotating LISA spacecrafts. The LISA Science Team has developed a Mock LISA Data Challenge intended to promote the testing of complicated nested search algorithms to detect the 100-1 millihertz frequency signals at amplitudes of 10E-21. However, it has become clear that, sequential search of the parameters is very time consuming and ultra-sensitive; hence, a new strategy has been developed. Parallelization of existing sequential search algorithms of Gravitational Wave signal identification consists of decomposing sequential search loops, beginning with outermost loops and working inward. In this process, the main challenge is to detect interdependencies among loops and partitioning the loops so as to preserve concurrency. Existing parallel programs are based upon either shared memory or distributed memory paradigms. In PVM, master and node programs are used to execute parallelization and process spawning. The PVM can handle process management and process addressing schemes using a virtual machine configuration. The task scheduling and the messaging and signaling can be implemented efficiently for the LISA Gravitational Wave search process using a master and 6 nodes. This approach is accomplished using a server that is available at NASA Ames Research Center, and has been dedicated to the LISA Data Challenge Competition. Historically, gravitational wave and source identification parameters have taken around 7 days in this dedicated single thread Linux based server. Using PVM approach, the parameter extraction problem can be reduced to within a day. The low frequency computation and a proxy signal-to-noise ratio are calculated in separate nodes that are controlled by the master using message and vector of data passing. The message passing among nodes follows a pattern of synchronous and asynchronous send-and-receive protocols. The communication model and the message buffers are allocated dynamically to address rapid search of gravitational wave source information in the Mock LISA data sets.

Thirumalainambi, Rajkumar↗

An Extension of the Time-Spectral Method to Overset Solvers

Relative motion in the Cartesian or overset framework causes certain spatial nodes to move in and out of the physical domain as they are dynamically blanked by moving solid bodies. This poses a problem for the conventional Time-Spectral approach, which expands the solution at every spatial node into a Fourier series spanning the period of motion. The proposed extension to the Time-Spectral method treats unblanked nodes in the conventional manner but expands the solution at dynamically blanked nodes in a basis of barycentric rational polynomials spanning partitions of contiguously defined temporal intervals. Rational polynomials avoid Runge's phenomenon on the equidistant time samples of these sub-periodic intervals. Fourier- and rational polynomial-based differentiation operators are used in tandem to provide a consistent hybrid Time-Spectral overset scheme capable of handling relative motion. The hybrid scheme is tested with a linear model problem and implemented within NASA's OVERFLOW Reynolds-averaged Navier- Stokes (RANS) solver. The hybrid Time-Spectral solver is then applied to inviscid and turbulent RANS cases of plunging and pitching airfoils and compared to time-accurate and experimental data. A limiter was applied in the turbulent case to avoid undershoots in the undamped turbulent eddy viscosity while maintaining accuracy. The hybrid scheme matches the performance of the conventional Time-Spectral method and converges to the time-accurate results with increased temporal resolution.

Leffell, Joshua Isaac↗

Computation of the inviscid supersonic flow about cones at large angles of attack by a floating discontinuity approach

The technique of floating shock fitting is adapted to the computation of the inviscid flowfield about circular cones in a supersonic free stream at angles of attack that exceed the cone half-angle. The resulting equations are applicable over the complete range of free-stream Mach numbers, angles of attack and cone half-angles for which the bow shock is attached. A finite difference algorithm is used to obtain the solution by an unsteady relaxation approach. The bow shock, embedded cross-flow shock, and vortical singularity in the leeward symmetry plane are treated as floating discontinuities in a fixed computational mesh. Where possible, the flowfield is partitioned into windward, shoulder, and leeward regions with each region computed separately to achieve maximum computational efficiency. An alternative shock fitting technique which treats the bow shock as a computational boundary is developed and compared with the floating-fitting approach. Several surface boundary condition schemes are also analyzed.

Daywitt, J.↗

Effects of partitioning and scheduling sparse matrix factorization on communication and load balance

A block based, automatic partitioning and scheduling methodology is presented for sparse matrix factorization on distributed memory systems. Using experimental results, this technique is analyzed for communication and load imbalance overhead. To study the performance effects, these overheads were compared with those obtained from a straightforward 'wrap mapped' column assignment scheme. All experimental results were obtained using test sparse matrices from the Harwell-Boeing data set. The results show that there is a communication and load balance tradeoff. The block based method results in lower communication cost whereas the wrap mapped scheme gives better load balance.

Venugopal, Sesh↗

Hierarchical Parallelism in Finite Difference Analysis of Heat Conduction

Based on the concept of hierarchical parallelism, this research effort resulted in highly efficient parallel solution strategies for very large scale heat conduction problems. Overall, the method of hierarchical parallelism involves the partitioning of thermal models into several substructured levels wherein an optimal balance into various associated bandwidths is achieved. The details are described in this report. Overall, the report is organized into two parts. Part 1 describes the parallel modelling methodology and associated multilevel direct, iterative and mixed solution schemes. Part 2 establishes both the formal and computational properties of the scheme.

Padovan, Joseph↗

E-beam generated holographic masks for optical vector-matrix multiplication

An optical vector matrix multiplication scheme that encodes the matrix elements as a holographic mask consisting of linear diffraction gratings is proposed. The binary, chrome on glass masks are fabricated by e-beam lithography. This approach results in a fairly simple optical system that promises both large numerical range and high accuracy. A partitioned computer generated hologram mask was fabricated and tested. This hologram was diagonally separated outputs, compact facets and symmetry about the axis. The resultant diffraction pattern at the output plane is shown. Since the grating fringes are written at 45 deg relative to the facet boundaries, the many on-axis sidelobes from each output are seen to be diagonally separated from the adjacent output signals.

Arnold, S. M.↗

Out-of-Core Streamline Visualization on Large Unstructured Meshes

It's advantageous for computational scientists to have the capability to perform interactive visualization on their desktop workstations. For data on large unstructured meshes, this capability is not generally available. In particular, particle tracing on unstructured grids can result in a high percentage of non-contiguous memory accesses and therefore may perform very poorly with virtual memory paging schemes. The alternative of visualizing a lower resolution of the data degrades the original high-resolution calculations. This paper presents an out-of-core approach for interactive streamline construction on large unstructured tetrahedral meshes containing millions of elements. The out-of-core algorithm uses an octree to partition and restructure the raw data into subsets stored into disk files for fast data retrieval. A memory management policy tailored to the streamline calculations is used such that during the streamline construction only a very small amount of data are brought into the main memory on demand. By carefully scheduling computation and data fetching, the overhead of reading data from the disk is significantly reduced and good memory performance results. This out-of-core algorithm makes possible interactive streamline visualization of large unstructured-grid data sets on a single mid-range workstation with relatively low main-memory capacity: 5-20 megabytes. Our test results also show that this approach is much more efficient than relying on virtual memory and operating system's paging algorithms.

Ueng, Shyh-Kuang↗

Trellis-coded multidimensional phase modulation

A 2L-dimensional multiple phase-shift keyed (MPSK) (L x MPSK) signal set is obtained by forming the Cartesian product of L two-dimensional MPSK signal sets. A systematic approach to partitioning L x MPSK signal sets that is based on block coding is used. An encoder system approach is developed. It incorporates the design of a differential precoder, a systematic convolutional encoder, and a signal set mapper. Trellis-coded L x 4PSK, L x 8PSK, and L x 16PSK modulation schemes are found for L = 1-4 and a variety of code rates and decoder complexities, many of which are fully transparent to discrete phase rotations of the signal set. The new codes achieve asymptotic coding gains up to 5.85 dB.

Pietrobon, Steven S.↗

Hierarchial implicit dynamic least-square solution algorithm

This paper develops an implicit type transient solution strategy which possesses hierarchial levels of application. In particular, due to the manner of formulation, stiffness updating, assembly inversion, solution constraint, as well as iteration are all performed at a localized level. The level of iterative calculations depends on the type of hierarchial partitioning employed, namely degree of freedom, nodal, elemental, material/nonlinear group, substructural, and so on. Since the iterative solution process and application of constraints are applied at a local level, the resulting so-called hierarchial implicit solution algorithm possesses very stable and efficient numerical properties and is highly storage efficient. To demonstrate the scheme, the results of several benchmark examples are presented. These enable comparisons with the Newton-Raphson solved implicit transient solution method. Overall the comparisons illustrate the superior stability and efficiency of the hierarchial scheme.

Padovan, J.↗

Zonal multigrid solution of compressible flow problems on unstructured and adaptive meshes

The simultaneous use of adaptive meshing techniques with a multigrid strategy for solving the 2-D Euler equations in the context of unstructured meshes is studied. To obtain optimal efficiency, methods capable of computing locally improved solutions without recourse to global recalculations are pursued. A method for locally refining an existing unstructured mesh, without regenerating a new global mesh is employed, and the domain is automatically partitioned into refined and unrefined regions. Two multigrid strategies are developed. In the first, time-stepping is performed on a global fine mesh covering the entire domain, and convergence acceleration is achieved through the use of zonal coarse grid accelerator meshes, which lie under the adaptively refined regions of the global fine mesh. Both schemes are shown to produce similar convergence rates to each other, and also with respect to a previously developed global multigrid algorithm, which performs time-stepping throughout the entire domain, on each mesh level. However, the present schemes exhibit higher computational efficiency due to the smaller number of operations on each level.

Mavriplis, Dimitri J.↗

Multiple-scale turbulence closure modeling of confined recirculating flows

A multiple-scale turbulence closure scheme is developed for the numerical predictions of confined recirculating flows. This model is based on the multiple-time-scale concepts of Hanjalic et al. (1980) and takes into account the non-equilibrium spectra energy transfer mechanism. Problems concerning new formulation of energy transfer rate equations and subsequent model coefficient redefinition and energy spectrum partition are discussed. Comparisons are made with several experiments of internal recirculating flows for the purpose of model validation. Numerical results using the present model show significant improvement of predictive capability over that obtained with the single-scale k-epsilon model and show promising potential for complex turbulent flow predictions.

Chen, C. P.↗

Microphysical Modelling of Polar Stratospheric Clouds During the 1999-2000 Winter

The evolution of the 1999-2000 Arctic winter has been examined using a microphysical/photochemical model run along diabatic trajectories. A large number of trajectories have been generated, filling the vortex throughout the region of polar stratospheric cloud (PSC) formation, and extending from November until the vortex breakup, in order to provide representative sampling of the evolution of PSCs and their effect on stratospheric chemistry. The 1999-2000 winter was particularly cold, allowing extensive PSC formation. Many trajectories have ten-day periods continuously below the Type I PSC threshold; significant periods of Type II PSCs are also indicated. The model has been used to test the extent and severity of denitrification and dehydration predicted using a range of different microphysical schemes. Scenarios in which freezing only occurs below the ice frost point (causing explicit coupling of denitrification and dehydration) have been tested, as well as scenarios with partial freezing at warmer temperatures (in which denitrification can occur independently of dehydration). The sensitivity to parameters such as aerosol freezing rates and heterogeneous freezing have been explored. Several scenarios cause sufficient denitrification to affect chlorine partitioning, and in turn, model-predicted ozone depletion, demonstrating that an improved understanding of the microphysics responsible for denitrification is necessary for understanding ozone loss rates.

Drdla, Katja↗

SEU System Analysis: Not Just the Sum of All Parts

Single event upset (SEU) analysis of complex systems is challenging. Currently, system SEU analysis is performed by component level partitioning and then either: the most dominant SEU cross-sections (SEUs) are used in system error rate calculations; or the partition SEUs are summed to eventually obtain a system error rate. In many cases, system error rates are overestimated because these methods generally overlook system level derating factors. The problem with overestimating is that it can cause overdesign and consequently negatively affect the following: cost, schedule, functionality, and validation/verification. The scope of this presentation is to discuss the risks involved with our current scheme of SEU analysis for complex systems; and to provide alternative methods for improvement.

Single Event Upset (SEU) Testing↗

Using SpF to Achieve Petascale for Legacy Pseudospectral Applications

Pseudospectral (PS) methods possess a number of characteristics (e.g., efficiency, accuracy, natural boundary conditions) that are extremely desirable for dynamo models. Unfortunately, dynamo models based upon PS methods face a number of daunting challenges, which include exposing additional parallelism, leveraging hardware accelerators, exploiting hybrid parallelism, and improving the scalability of global memory transposes. Although these issues are a concern for most models, solutions for PS methods tend to require far more pervasive changes to underlying data and control structures. Further, improvements in performance in one model are difficult to transfer to other models, resulting in significant duplication of effort across the research community. We have developed an extensible software framework for pseudospectral methods called SpF that is intended to enable extreme scalability and optimal performance. Highlevel abstractions provided by SpF unburden applications of the responsibility of managing domain decomposition and load balance while reducing the changes in code required to adapt to new computing architectures. The key design concept in SpF is that each phase of the numerical calculation is partitioned into disjoint numerical kernels that can be performed entirely inprocessor. The granularity of domain decomposition provided by SpF is only constrained by the datalocality requirements of these kernels. SpF builds on top of optimized vendor libraries for common numerical operations such as transforms, matrix solvers, etc., but can also be configured to use open source alternatives for portability. SpF includes several alternative schemes for global data redistribution and is expected to serve as an ideal testbed for further research into optimal approaches for different network architectures. In this presentation, we will describe our experience in porting legacy pseudospectral models, MoSST and DYNAMO, to use SpF as well as present preliminary performance results provided by the improved scalability.

DYNAMO↗

Galileo mission planning for Low Gain Antenna based operations

The Galileo mission operations concept is undergoing substantial redesign, necessitated by the deployment failure of the High Gain Antenna, while the spacecraft is on its way to Jupiter. The new design applies state-of-the-art technology and processes to increase the telemetry rate available through the Low Gain Antenna and to increase the information density of the telemetry. This paper describes the mission planning process being developed as part of this redesign. Principal topics include a brief description of the new mission concept and anticipated science return (these have been covered more extensively in earlier papers), identification of key drivers on the mission planning process, a description of the process and its implementation schedule, a discussion of the application of automated mission planning tool to the process, and a status report on mission planning work to date. Galileo enhancements include extensive reprogramming of on-board computers and substantial hard ware and software upgrades for the Deep Space Network (DSN). The principal mode of operation will be onboard recording of science data followed by extended playback periods. A variety of techniques will be used to compress and edit the data both before recording and during playback. A highly-compressed real-time science data stream will also be important. The telemetry rate will be increased using advanced coding techniques and advanced receivers. Galileo mission planning for orbital operations now involves partitioning of several scarce resources. Particularly difficult are division of the telemetry among the many users (eleven instruments, radio science, engineering monitoring, and navigation) and allocation of space on the tape recorder at each of the ten satellite encounters. The planning process is complicated by uncertainty in forecast performance of the DSN modifications and the non-deterministic nature of the new data compression schemes. Key mission planning steps include quantifying resource or capabilities to be allocated, prioritizing science observations and estimating resource needs for each, working inter-and intra-orbit trades of these resources among the Project elements, and planning real-time science activity. The first major mission planning activity, a high level, orbit-by-orbit allocation of resources among science objectives, has already been completed; and results are illustrated in the paper. To make efficient use of limited resources, Galileo mission planning will rely on automated mission planning tools capable of dealing with interactions among time-varying downlink capability, real-time science and engineering data transmission, and playback of recorded data. A new generic mission planning tool is being adapted for this purpose.

Gershman, R.↗

The Perils of Partition: Difficulties in Retrieving Magma Compositions from Chemically Equilibrated Basaltic Meteorites

The chemical compositions of magmas can be derived from the compositions of their equilibrium minerals through mineral/magma partition coefficients. This method cannot be applied safely to basaltic rocks, either solidified lavas or cumulates, which have chemically equilibrated or partially equilibrated at subsolidus temperatures, i.e., in the absence of magma. Applying mineral/ melt partition coefficients to mineral compositions from such rocks will typically yield 'magma compositions' that are strongly fractionated and unreasonably enriched in incompatible elements (e.g., REE's). In the absence of magma, incompatible elements must go somewhere; they are forced into minerals (e.g., pyroxenes, plagioclase) at abundance levels far beyond those established during normal mineral/magma equilibria. Further, using mineral/magma partition coefficients with such rocks may suggest that different minerals equilibrated with different magmas, and the fractionation sequence of those melts (i.e., enrichment in incompatible elements) may not be consistent with independent constraints on the order of crystallization. Subsolidus equilibration is a reasonable cause for incompatible- element-enriched minerals in some eucrites, diogenites, and martian meteorites and offers a simple alternative to petrogenetic schemes involving highly fractionated magmas or magma infiltration metasomatism.

Treiman, Allan H.↗

LSI computer fabrication - SUMC/DV.

The SUMC/DV (Space Ultrareliable Modular Computer Demonstration Vehicle), which required the design and fabrication of ten different array types, was designed, partitioned, assembled, tested, and made completely operational in less than one year. This paper describes the assembly, the physical and electrical characteristics, and the basic electrical testing procedures used on SUMC/DV. The paper includes a description of the packaging concepts, the physical assembly, the fabrication and testing of the various components including the LSI/CMOS arrays, the hierarchy of the memory complement, the clock generation and distribution scheme, additional system testing techniques, and the procedure employed in the electrical checkout of the system.

Feller, A.↗

Optimal and Local Connectivity Between Neuron and Synapse Array in the Quantum Dot/Silicon Brain

This innovation is used to connect between synapse and neuron arrays using nanowire in quantum dot and metal in CMOS (complementary metal oxide semiconductor) technology to enable the density of a brain-like connection in hardware. The hardware implementation combines three technologies: 1. Quantum dot and nanowire-based compact synaptic cell (50x50 sq nm) with inherently low parasitic capacitance (hence, low dynamic power approx.l0(exp -11) watts/synapse), 2. Neuron and learning circuits implemented in 50-nm CMOS technology, to be integrated with quantum dot and nanowire synapse, and 3. 3D stacking approach to achieve the overall numbers of high density O(10(exp 12)) synapses and O(10(exp 8)) neurons in the overall system. In a 1-sq cm of quantum dot layer sitting on a 50-nm CMOS layer, innovators were able to pack a 10(exp 6)-neuron and 10(exp 10)-synapse array; however, the constraint for the connection scheme is that each neuron will receive a non-identical 10(exp 4)-synapse set, including itself, via its efficacy of the connection. This is not a fully connected system where the 100x100 synapse array only has a 100-input data bus and 100-output data bus. Due to the data bus sharing, it poses a great challenge to have a complete connected system, and its constraint within the quantum dot and silicon wafer layer. For an effective connection scheme, there are three conditions to be met: 1. Local connection. 2. The nanowire should be connected locally, not globally from which it helps to maximize the data flow by sharing the same wire space location. 3. Each synapse can have an alternate summation line if needed (this option is doable based on the simple mask creation). The 10(exp 3)x10(exp 3)-neuron array was partitioned into a 10-block, 10(exp 2)x10(exp 3)-neuron array. This building block can be completely mapped within itself (10,000 synapses to a neuron).

Duong, Tuan A.↗