Search NASASearch

Engineering topics

Chancellor, Marisa K.

Publications and source records attributed to Chancellor, Marisa K..

At least 19 records

UFLIC: A Line Integral Convolution Algorithm for Visualizing Unsteady Flows

This paper presents an algorithm, UFLIC (Unsteady Flow LIC), to visualize vector data in unsteady flow fields. Using the Line Integral Convolution (LIC) as the underlying method, a new convolution algorithm is proposed that can effectively trace the flow's global features over time. The new algorithm consists of a time-accurate value depositing scheme and a successive feed-forward method. The value depositing scheme accurately models the flow advection, and the successive feed-forward method maintains the coherence between animation frames. Our new algorithm can produce time-accurate, highly coherent flow animations to highlight global features in unsteady flow fields. CFD scientists, for the first time, are able to visualize unsteady surface flows using our algorithm.

Shen, Han-Wei

Improved Boundary Conditions for Cell-centered Difference Schemes

Cell-centered finite-volume (CCFV) schemes have certain attractive properties for the solution of the equations governing compressible fluid flow. Among others, they provide a natural vehicle for specifying flux conditions at the boundaries of the physical domain. Unfortunately, they lead to slow convergence for numerical programs utilizing them. In this report a method for investigating and improving the convergence of CCFV schemes is presented, which focuses on the effect of the numerical boundary conditions. The key to the method is the computation of the spectral radius of the iteration matrix of the entire demoralized system of equations, not just of the interior point scheme or the boundary conditions.

VanderWijngaart, Rob F.

Efficacy of Code Optimization on Cache-based Processors

The current common wisdom in the U.S. is that the powerful, cost-effective supercomputers of tomorrow will be based on commodity (RISC) micro-processors with cache memories. Already, most distributed systems in the world use such hardware as building blocks. This shift away from vector supercomputers and towards cache-based systems has brought about a change in programming paradigm, even when ignoring issues of parallelism. Vector machines require inner-loop independence and regular, non-pathological memory strides (usually this means: non-power-of-two strides) to allow efficient vectorization of array operations. Cache-based systems require spatial and temporal locality of data, so that data once read from main memory and stored in high-speed cache memory is used optimally before being written back to main memory. This means that the most cache-friendly array operations are those that feature zero or unit stride, so that each unit of data read from main memory (a cache line) contains information for the next iteration in the loop. Moreover, loops ought to be 'fat', meaning that as many operations as possible are performed on cache data-provided instruction caches do not overflow and enough registers are available. If unit stride is not possible, for example because of some data dependency, then care must be taken to avoid pathological strides, just ads on vector computers. For cache-based systems the issues are more complex, due to the effects of associativity and of non-unit block (cache line) size. But there is more to the story. Most modern micro-processors are superscalar, which means that they can issue several (arithmetic) instructions per clock cycle, provided that there are enough independent instructions in the loop body. This is another argument for providing fat loop bodies. With these restrictions, it appears fairly straightforward to produce code that will run efficiently on any cache-based system. It can be argued that although some of the important computational algorithms employed at NASA Ames require different programming styles on vector machines and cache-based machines, respectively, neither architecture class appeared to be favored by particular algorithms in principle. Practice tells us that the situation is more complicated. This report presents observations and some analysis of performance tuning for cache-based systems. We point out several counterintuitive results that serve as a cautionary reminder that memory accesses are not the only factors that determine performance, and that within the class of cache-based systems, significant differences exist.

VanderWijngaart, Rob F.

The Need for Vendor Source Code at NAS

The Numerical Aerodynamic Simulation (NAS) Facility has a long standing practice of maintaining buildable source code for installed hardware. There are two reasons for this: NAS's designated pathfinding role, and the need to maintain a smoothly running operational capacity given the widely diversified nature of the vendor installations. NAS has a need to maintain support capabilities when vendors are not able; diagnose and remedy hardware or software problems where applicable; and to support ongoing system software development activities whether or not the relevant vendors feel support is justified. This note provides an informal history of these activities at NAS, and brings together the general principles that drive the requirement that systems integrated into the NAS environment run binaries built from source code, onsite.

Carter, Russell

Foundations for Measuring Volume Rendering Quality

The goal of this paper is to provide a foundation for objectively comparing volume rendered images. The key elements of the foundation are: (1) a rigorous specification of all the parameters that need to be specified to define the conditions under which a volume rendered image is generated; (2) a methodology for difference classification, including a suite of functions or metrics to quantify and classify the difference between two volume rendered images that will support an analysis of the relative importance of particular differences. The results of this method can be used to study the changes caused by modifying particular parameter values, to compare and quantify changes between images of similar data sets rendered in the same way, and even to detect errors in the design, implementation or modification of a volume rendering system. If one has a benchmark image, for example one created by a high accuracy volume rendering system, the method can be used to evaluate the accuracy of a given image.

Williams, Peter L.

Onward to Petaflops Computing

With programs such as the US High Performance Computing and Communications Program (HPCCP), the attention of scientists and engineers worldwide has been focused on the potential of very high performance scientific computing, namely systems that are hundreds or thousands of times more powerful than those typically available in desktop systems at any given point in time. Extending the frontiers of computing in this manner has resulted in remarkable advances, both in computing technology itself and also in the various scientific and engineering disciplines that utilize these systems. Within the month or two, a sustained rate of 1 Tflop/s (also written 1 teraflops, or 10(exp 12) floating-point operations per second) is likely to be achieved by the 'ASCI Red' system at Sandia National Laboratory in New Mexico. With this objective in sight, it is reasonable to ask what lies ahead for high-end computing.

Bailey, David H.

A Portable MPI Implementation of the SPAI Preconditioner in ISIS++

A parallel MPI implementation of the Sparse Approximate Inverse (SPAI) preconditioner is described. SPAI has proven to be a highly effective preconditioner, and is inherently parallel because it computes columns (or rows) of the preconditioning matrix independently. However, there are several problems that must be addressed for an efficient MPI implementation: load balance, latency hiding, and the need for one-sided communication. The effectiveness, efficiency, and scaling behavior of our implementation will be shown for different platforms.

Barnard, Stephen T.

NAS Parallel Benchmark Results 11-96

The NAS Parallel Benchmarks have been developed at NASA Ames Research Center to study the performance of parallel supercomputers. The eight benchmark problems are specified in a "pencil and paper" fashion. In other words, the complete details of the problem to be solved are given in a technical document, and except for a few restrictions, benchmarkers are free to select the language constructs and implementation techniques best suited for a particular system. These results represent the best results that have been reported to us by the vendors for the specific 3 systems listed. In this report, we present new NPB (Version 1.0) performance results for the following systems: DEC Alpha Server 8400 5/440, Fujitsu VPP Series (VX, VPP300, and VPP700), HP/Convex Exemplar SPP2000, IBM RS/6000 SP P2SC node (120 MHz), NEC SX-4/32, SGI/CRAY T3E, SGI Origin200, and SGI Origin2000. We also report High Performance Fortran (HPF) based NPB results for IBM SP2 Wide Nodes, HP/Convex Exemplar SPP2000, and SGI/CRAY T3D. These results have been submitted by Applied Parallel Research (APR) and Portland Group Inc. (PGI). We also present sustained performance per dollar for Class B LU, SP and BT benchmarks.

Bailey, David H.

Molecular Dynamics Simulation of a Multi-Walled Carbon Nanotube Based Gear

We used molecular dynamics to investigate the properties of a multi-walled carbon nanotube based gear. Previous work computationally suggested that molecular gears fashioned from (14,0) single-walled carbon nanotubes operate well at 50-100 gigahertz. The gears were formed from nanotubes with teeth added via a benzyne reaction known to occur with C60. A modified, parallelized version of Brenner's potential was used to model interatomic forces within each molecule. A Leonard-Jones 6-12 potential was used for forces between molecules. The gear in this study was based on the smallest multi-walled nanotube supported by some experimental evidence. Each gear was a (52,0) nanotube surrounding a (37,10) nanotube with approximate 20.4 and 16,8 A radii respectively. These sizes were chosen to be consistent with inter-tube spacing observed by and were slightly larger than graphite inter-layer spacings. The benzyne teeth were attached via 2+4 cycloaddition to exterior of the (52,0) tube. 2+4 bonds were used rather than the 2+2 bonds observed by Hoke since 2+4 bonds are preferred by naphthalene and quantum calculations by Jaffe suggest that 2+4 bonds are preferred on carbon nanotubes of sufficient diameter. One gear was 'powered' by forcing the atoms near the end of the outside buckytube to rotate to simulate a motor. A second gear was allowed to rotate by keeping the atoms near the end of its outside buckytube on a cylinder. The ends of both gears were constrained to stay in an approximately constant position relative to each other, simulating a casing, to insure that the gear teeth meshed. The stiff meshing aromatic gear teeth transferred angular momentum from the powered gear to the driven gear. The simulation was performed in a vacuum and with a software thermostat. Preliminary results suggest that the powered gear had trouble turning the driven gear without slip. The larger radius and greater mass of these gears relative to the (14,0) gears previously studied requires a smaller rotation rate and multiple rows of teeth to avoid excessive force on the gear teeth resulting, in slip and failure of the driven gear to turn. We hope that studies such as these will eventually lead to synthesis of components that can be assembled into atomically precise fullerene machines. These machines, in turn, may someday be used in machine-phase fullerene materials with remarkable properties.

Han, Jie

Machine Phase Fullerene Nanotechnology: 1996

NASA has used exotic materials for spacecraft and experimental aircraft to good effect for many decades. In spite of many advances, transportation to space still costs about $10,000 per pound. Drexler has proposed a hypothetical nanotechnology based on diamond and investigated the properties of such molecular systems. These studies and others suggest enormous potential for aerospace systems. Unfortunately, methods to realize diamonoid nanotechnology are at best highly speculative. Recent computational efforts at NASA Ames Research Center and computation and experiment elsewhere suggest that a nanotechnology of machine phase functionalized fullerenes may be synthetically relatively accessible and of great aerospace interest. Machine phase materials are (hypothetical) materials consisting entirely or in large part of microscopic machines. In a sense, most living matter fits this definition. To begin investigation of fullerene nanotechnology, we used molecular dynamics to study the properties of carbon nanotube based gears and gear/shaft configurations. Experiments on C60 and quantum calculations suggest that benzyne may react with carbon nanotubes to form gear teeth. Han has computationally demonstrated that molecular gears fashioned from (14,0) single-walled carbon nanotubes and benzyne teeth should operate well at 50-100 gigahertz. Results suggest that rotation can be converted to rotating or linear motion, and linear motion may be converted into rotation. Preliminary results suggest that these mechanical systems can be cooled by a helium atmosphere. Furthermore, Deepak has successfully simulated using helical electric fields generated by a laser to power fullerene gears once a positive and negative charge have been added to form a dipole. Even with mechanical motion, cooling, and power; creating a viable nanotechnology requires support structures, computer control, a system architecture, a variety of components, and some approach to manufacture. Additional information is contained within the original extended abstract.

Globus, Al

Formation of Carbon Nanotube Based Gears: Quantum Chemistry and Molecular Mechanics Study of the Electrophilic Addition of o-Benzyne to Fullerenes, Graphene, and Nanotubes

Considerable progress has been made in recent years in chemical functionalization of fullerene molecules. In some cases, the predominant reaction products are different from those obtained (using the same reactants) from polycyclic aromatic hydrocarbons (PAHs). One such example is the cycloaddition of o-benzyne to C60. It is well established that benzyne adds across one of the rings in naphthalene, anthracene and other PAHs forming the [2+4] cycloaddition product (benzobicyclo[2.2.2.]-octatriene with naphthalene and triptycene with anthracene). However, Hoke et al demonstrated that the only reaction path for o-benzyne with C60 leads to the [2+2] cycloaddition product in which benzyne adds across one of the interpentagonal bonds (forming a cyclobutene ring in the process). Either reaction product results in a loss of aromaticity and distortion of the PAH or fullerene substrate, and in a loss of strain in the benzyne. It is not clear, however, why different products are preferred in these cases. In the current paper, we consider the stability of benzyne-nanotube adducts and the ability of Brenner's potential energy model to describe the structure and stability of these adducts. The Brenner potential has been widely used for describing diamondoid and graphitic carbon. Recently it has also been used for molecular mechanics and molecular dynamics simulations of fullerenes and nanotubes. However, it has not been tested for the case of functionalized fullerenes (especially with highly strained geometries). We use the Brenner potential for our companion nanogear simulations and believe that it should be calibrated to insure that those simulations are physically reasonable. In the present work, Density Functional theory (DFT) calculations are used to determine the preferred geometric structures and energetics for this calibration. The DFT method is a kind of ab initio quantum chemistry method for determining the electronic structure of molecules. For a given basis set expansion, it is comparable in accuracy to the MP2 method (better than Hartree Fock, but less accurate than more extensive electron correlation methods such as MP4 or CCSD). However, for systems with large numbers of basis functions it more efficient than any other methods that include electron correlation effects. In this presentation we show the results of DFT calculations for the reaction of benzyne with naphthalene, C60, and nanotube models. We compare energies for [2+2] and [2+4] cycloaddition products. The preferred products for the naphthalene and C60 reactions have been determined by experiment and, thus, these cases serve as a validation of our quantum chemical approach. We also compare the DFT and Brenner potential results. Finally we can predict the likelihood of reaction between benzyne and nanotubes.

Jaffe, Richard

Petaflops Computing: The Key Algorithmic Challenges

The prospect of petaflops-class computers brings to the fore some important algorithmic issues that have been considered in the high performance computing community for several years. Key among them are (1) concurrency (whether the fundamental concurrency of an algorithm is sufficient to keep thousands of processors productively busy); (2) data locality; (3) latency tolerance; and (4) memory and operation count scaling. This introductory presentation will give an overview of these issues.

Bailey, David H.

Molecular Nanotechnology and Designs of Future

Reviewing the status of current approaches and future projections, as already published in the scientific journals and books, the talk will summarize the direction in which computational and experimental molecular nanotechnologies are progressing. Examples of nanotechnological approach to the concepts of design and simulation of atomically precise materials in a variety of interdisciplinary areas will be presented. The concepts of hypothetical molecular machines and assemblers as explained in Drexler's and Merckle's already published work and Han et. al's WWW distributed molecular gears will be explained.

Srivastava, Deepak

Molecular Dynamics Simulations of Laser Powered Carbon Nanotube Gears

Dynamics of laser powered carbon nanotube gears is investigated by molecular dynamics simulations with Brenner's hydrocarbon potential. We find that when the frequency of the laser electric field is much less than the intrinsic frequency of the carbon nanotube, the tube exhibits an oscillatory pendulam behavior. However, a unidirectional rotation of the gear with oscillating frequency is observed under conditions of resonance between the laser field and intrinsic gear frequencies. The operating conditions for stable rotations of the nanotube gears, powered by laser electric fields are explored, in these simulations.

Srivastava, Deepak

Bohm's Quantum Potential and the Visualization of Molecular Structure

David Bohm's ontological interpretation of quantum theory can shed light on otherwise counter-intuitive quantum mechanical phenomena including chemical bonding. In the field of quantum chemistry, Richard Bader has shown that the topology of the Laplacian of the electronic charge density characterizes many features of molecular structure and reactivity. Visual and computational examination suggests that the Laplacian of Bader and the quantum potential of Bohm are morphologically equivalent. It appears that Bohmian mechanics and the quantum potential can make chemistry as clear as they makes physics.

Levit, Creon

Communication Studies of DMP and SMP Machines

Understanding the interplay between machines and problems is key to obtaining high performance on parallel machines. This paper investigates the interplay between programming paradigms and communication capabilities of parallel machines. In particular, we explicate the communication capabilities of the IBM SP-2 distributed-memory multiprocessor and the SGI PowerCHALLENGEarray symmetric multiprocessor. Two benchmark problems of bitonic sorting and Fast Fourier Transform are selected for experiments. Communication-efficient algorithms are developed to exploit the overlapping capabilities of the machines. Programs are written in Message-Passing Interface for portability and identical codes are used for both machines. Various data sizes and message sizes are used to test the machines' communication capabilities. Experimental results indicate that the communication performance of the multiprocessors are consistent with the size of messages. The SP-2 is sensitive to message size but yields a much higher communication overlapping because of the communication co-processor. The PowerCHALLENGEarray is not highly sensitive to message size and yields a low communication overlapping. Bitonic sorting yields lower performance compared to FFT due to a smaller computation-to-communication ratio.

Sohn, Andrew

Onward to Petaflops Computing

With the recent demonstration of a computing rate of one Tflop/s at Sandia National Lab, one might ask what lies ahead for high-end computing. The next major milestone is a sustained rate of one Pflop/s (also written one petaflops, or 10(exp 15) floating-point operations per second). It should be emphasized that we could just as well use the term "peta-ops", since it appears that large scientific systems will be required to perform intensive integer and logical computation in addition to floating-point operations, and completely non- floating-point applications are likely to be important as well. In addition to prodigiously high computational performance, such systems must of necessity feature very large main memories, between ten Tbyte (10(exp 13) byte) and one Pbyte (10 (exp 15) byte) depending on application, as well as commensurate I/O bandwidth and huge mass storage facilities. The current consensus of scientists who have performed initial studies in this field is that "affordable" petaflops systems may be feasible by the year 2010, assuming that certain key technologies continue to progress at current rates. A sustained petaflops computing capability however is a daunting challenge; it appears significantly more challenging from today's state-of-the-art than achieving one Tflop/s has been from the level of one Gflop/s about 12 years ago. Challenges are faced in the arena of device technology, system architecture, system software, algorithms and applications. This talk will give an overview of some of these challenges, and describe some of the recent initiatives to address them.

Bailey, David H.

Toroidal Single Wall Carbon Nanotubes in Fullerene Crop Circles

We investigate energetics and structure of circular and polygonal single wall carbon nanotubes (SWNTs) using large scale molecular simulations on NAS SP2, motivated by their unusual electronic and magnetic properties. The circular tori are formed by bending tube (no net whereas the polygonal tori are constructed by turning the joint of two tubes of (n, n), (n+1, n-1) and (n+2, n-2) with topological pentagon-heptagon defect, in which n =5, 8 and 10. The strain energy of circular tori relative to straight tube decreases by I/D(sup 2) where D is torus diameter. As D increases, these tori change from buckling to an energetically stable state. The stable tori are perfect circular in both toroidal and tubular geometry with strain less than 0. 03 eV/atom when D greater than 10, 20 and 40 nm for torus (5,5), (8,8) and (10, 10). Polygonal tori, whose strain is proportional to the number of defects and I/D are energetically stable even for D less than 10 nm. However, their strain is higher than that of perfect circular tori. In addition, the local maximum strain of polygonal tori is much higher than that of perfect circular tori. It is approx. 0.03 eV/atom or less for perfect circular torus (5,5), but 0.13 and 0.21 eV/atom for polygonal tori (6,4)/(5,5) and (7,3)/(5,5). Therefore, we conclude that the circular tori with no topological defects are more energetically stable and kinetically accessible than the polygonal tori containing the pentagon-heptagon defects for the laser-grown SWNTs and Fullerene crop circles.

Han, Jie