Search NASA⌕ Search

SEARCH · Search NASA

Results for “Kernel”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

Robust Multigrid Smoothers for Three Dimensional Elliptic Equations with Strong Anisotropies

We discuss the behavior of several plane relaxation methods as multigrid smoothers for the solution of a discrete anisotropic elliptic model problem on cell-centered grids. The methods compared are plane Jacobi with damping, plane Jacobi with partial damping, plane Gauss-Seidel, plane zebra Gauss-Seidel, and line Gauss-Seidel. Based on numerical experiments and local mode analysis, we compare the smoothing factor of the different methods in the presence of strong anisotropies. A four-color Gauss-Seidel method is found to have the best numerical and architectural properties of the methods considered in the present work. Although alternating direction plane relaxation schemes are simpler and more robust than other approaches, they are not currently used in industrial and production codes because they require the solution of a two-dimensional problem for each plane in each direction. We verify the theoretical predictions of Thole and Trottenberg that an exact solution of each plane is not necessary and that a single two-dimensional multigrid cycle gives the same result as an exact solution, in much less execution time. Parallelization of the two-dimensional multigrid cycles, the kernel of the three-dimensional implicit solver, is also discussed. Alternating-plane smoothers are found to be highly efficient multigrid smoothers for anisotropic elliptic problems.

Llorente, Ignacio M.↗

CAPRI (Computational Analysis PRogramming Interface): A Solid Modeling Based Infra-Structure for Engineering Analysis and Design Simulations

CAPRI is a CAD-vendor neutral application programming interface designed for the construction of analysis and design systems. By allowing access to the geometry from within all modules (grid generators, solvers and post-processors) such tasks as meshing on the actual surfaces, node enrichment by solvers and defining which mesh faces are boundaries (for the solver and visualization system) become simpler. The overall reliance on file 'standards' is minimized. This 'Geometry Centric' approach makes multi-physics (multi-disciplinary) analysis codes much easier to build. By using the shared (coupled) surface as the foundation, CAPRI provides a single call to interpolate grid-node based data from the surface discretization in one volume to another. Finally, design systems are possible where the results can be brought back into the CAD system (and therefore manufactured) because all geometry construction and modification are performed using the CAD system's geometry kernel.

Haimes, Robert↗

Nonlinear Rescaling and Proximal-Like Methods in Convex Optimization

The nonlinear rescaling principle (NRP) consists of transforming the objective function and/or the constraints of a given constrained optimization problem into another problem which is equivalent to the original one in the sense that their optimal set of solutions coincides. A nonlinear transformation parameterized by a positive scalar parameter and based on a smooth scaling function is used to transform the constraints. The methods based on NRP consist of sequential unconstrained minimization of the classical Lagrangian for the equivalent problem, followed by an explicit formula updating the Lagrange multipliers. We first show that the NRP leads naturally to proximal methods with an entropy-like kernel, which is defined by the conjugate of the scaling function, and establish that the two methods are dually equivalent for convex constrained minimization problems. We then study the convergence properties of the nonlinear rescaling algorithm and the corresponding entropy-like proximal methods for convex constrained optimization problems. Special cases of the nonlinear resealing algorithm are presented. In particular a new class of exponential penalty-modified barrier functions methods is introduced.

Polyak, Roman↗

Object-Oriented Design for Sparse Direct Solvers

We discuss the object-oriented design of a software package for solving sparse, symmetric systems of equations (positive definite and indefinite) by direct methods. At the highest layers, we decouple data structure classes from algorithmic classes for flexibility. We describe the important structural and algorithmic classes in our design, and discuss the trade-offs we made for high performance. The kernels at the lower layers were optimized by hand. Our results show no performance loss from our object-oriented design, while providing flexibility, case of use, and extensibility over solvers using procedural design.

Dobrian, Florin↗

Analysis of an Interface Crack for a Functionally Graded Strip Sandwiched between Two Homogeneous Layers of Finite Thickness

The interface crack problem for a composite layer that consists of a homogeneous substrate, coating and a non-homogeneous interface was formulated for singular integral equations with Cauchy kernels and integrated using the Lobatto-Chebyshev collocation technique. Mixed-mode Stress Intensity Factors and Strain Energy Release Rates were calculated. The Stress Intensity Factors were compared for accuracy with relevant results previously published. The parametric studies were conducted for the various thickness of each layer and for various non-homogeneity ratios. Particular application to the Zirconia thermal barrier on steel substrate is demonstrated.

Shbeeh, N. I.↗

Improving the Accuracy of Quadrature Method Solutions of Fredholm Integral Equations That Arise from Nonlinear Two-Point Boundary Value Problems

In this paper we are concerned with high-accuracy quadrature method solutions of nonlinear Fredholm integral equations of the form y(x) = r(x) + definite integral of g(x, t)F(t,y(t))dt with limits between 0 and 1,0 less than or equal to x les than or equal to 1, where the kernel function g(x,t) is continuous, but its partial derivatives have finite jump discontinuities across x = t. Such integral equations arise, e.g., when one applied Green's function techniques to nonlinear two-point boundary value problems of the form y "(x) =f(x,y(x)), 0 less than or equal to x less than or equal to 1, with y(0) = y(sub 0) and y(l) = y(sub l), or other linear boundary conditions. A quadrature method that is especially suitable and that has been employed for such equations is one based on the trepezoidal rule that has a low accuracy. By analyzing the corresponding Euler-Maclaurin expansion, we derive suitable correction terms that we add to the trapezoidal rule, thus obtaining new numerical quadrature formulas of arbitrarily high accuracy that we also use in defining quadrature methods for the integral equations above. We prove an existence and uniqueness theorem for the quadrature method solutions, and show that their accuracy is the same as that of the underlying quadrature formula. The solution of the nonlinear systems resulting from the quadrature methods is achieved through successive approximations whose convergence is also proved. The results are demonstrated with numerical examples.

Sidi, Avram↗

Northern and Southern Hemisphere Ground-Based Infrared Spectroscopic Measurements of Tropospheric Carbon Monoxide and Ethane

Time series of CO and C2H, measurements have been derived from high-resolution infrared solar spectra recorded in Lauder, New Zealand (45.0 degrees S, 169.7 degrees E, altitude 0.37 km), and at the U.S. National Solar Observatory (31.9 degrees N, 11, 1.6 degrees W, altitude 2.09 km) on Kitt Peak. Lauder observations were obtained between July 1993 and November 1997, while the Kitt Peak measurements were recorded between May 1977 and December 1997. Both databases were analyzed with spectroscopic parameters that included significant improvements for C2H6 relative to previous studies. Target CO and C2H6 lines were selected to achieve similar vertical samplings based on averaging kernels. These calculations show that partial columns from layers extending from the surface to the mean tropopause and from the mean tropopause to 100 km are nearly independent. Retrievals based on a semiempirical application of the Rodgers optimal estimation technique are reported for the lower layer, which has a broad maximum in sensitivity in the upper troposphere. The Lauder CO and C2H, partial columns exhibit highly asymmetrical seasonal cycles with minima in austral autumn and sharp peaks in austral spring. The spring maxima are the result of tropical biomass burning emissions followed by deep convective vertical transport to the upper troposphere and long-range horizontal transport. Significant year-to-year variations are observed for both CO and C2H6, but the measured trends, (+0.37 +/- 0.57)% yr(exp -1) and (-0.64 +/- 0.79)% yr(exp -1), I sigma, respectively, indicate no significant long-term changes. The Kitt Peak data also exhibit CO and C2H6, seasonal variations in the lower layer with trends equal to (-0.27 +/- 0.17)% yr(exp -1) and (-1.20 +/- 0.35)% yr(exp -1), 1 sigma, respectively. Hence a decrease in the Kitt Peak tropospheric C2H6 column has been detected, though the CO trend is not significant. Both measurement sets are compared with previous observations, reported trends, and three-dimensional model calculations.

Rinsland, Curtis P.↗

Improving the Accuracy of Quadrature Method Solutions of Fredholm Integral Equations that Arise from Nonlinear Two-Point Boundary Value Problems

In this paper we are concerned with high-accuracy quadrature method solutions of nonlinear Fredholm integral equations of the form y(x) = r(x) + integral(0 to 1) g(x,t) F(t, y(t)) dt, 0 less than or equal to x less than or equal to 1, where the kernel function g(x,t) is continuous, but its partial derivatives have finite jump discontinuities across x = t. Such integrals equations arise, e.g., when one applies Green's function techniques to nonlinear two-point boundary value problems of the form U''(x) = f(x,y(x)), 0 less than or equal to x less than or equal to 1, with y(0) = y(sub 0) and g(l) = y(sub 1), or other linear boundary conditions. A quadrature method that is especially suitable and that has been employed for such equations is one based on the trapezoidal rule that has a low accuracy. By analyzing the corresponding Euler-Maclaurin expansion, we derive suitable correction terms that we add to the trapezoidal thus obtaining new numerical quadrature formulas of arbitrarily high accuracy that we also use in defining quadrature methods for the integral equations above. We prove an existence and uniqueness theorem for the quadrature method solutions, and show that their accuracy is the same as that of the underlying quadrature formula. The solution of the nonlinear systems resulting from the quadrature methods is achieved through successive approximations whose convergence is also proved. The results are demonstrated with numerical examples.

Sidi, Avram↗

Northern and Southern Hemisphere Ground-Based Infrared Spectroscopic Measurements of Tropospheric Carbon Monoxide and Ethane

Time series of CO and C2H6 measurements have been derived from high resolution infrared solar spectra recorded in Lauder, New Zealand (45.0 deg S, 169.7 deg E, altitude 0.37 km) and at the U. S. National Solar Observatory (31.90 deg N, 111.6 deg W, altitude 2.09 km) on Kitt Peak. Lauder observations were obtained between July 1993 and November 1997 while the Kitt Peak measurements were recorded between May 1977 and December 1997. Both databases were analyzed with spectroscopic parameters that included significant improvements for C2H6 relative to previous studies. Target CO and C2H6 lines were selected to achieve similar vertical samplings based on averaging kernels. These calculations show that partial columns from layers extending from the surface to the mean tropopause and from the mean tropopause to 100 km are nearly independent. Retrievals based on a semiempirical application of the Rodgers optimal estimation technique are reported for the lower layer, which has a broad maximum in sensitivity in the upper troposphere. The Lauder CO and C2H6 partial columns exhibit highly asymmetrical seasonal cycles with minima in austral autumn and sharp peaks in austral spring. The spring maxima are the result of tropical biomass burning emissions followed by deep convective vertical transport to the upper troposphere and long-range horizontal transport. Significant year-to-year variations are observed for both CO and C2H6, but the measured trends, (+0.37 +/- 0.57)%/ yr and (-0.64 +/- 0.79)%/ yr, 1 sigma, respectively, indicate no significant long-term changes. The Kitt Peak data also exhibit CO and C2H6 seasonal variations in the lower layer with trends equal to (-0.27 +/- 0.17)%/ yr and (-1.20 +/- 0.35)%/ yr, 1 sigma, respectively. Hence, a decrease in the Kitt Peak tropospheric C2H6 column has been detected, though the CO trend is not significant. Both measurement sets are compared with previous observations, reported trends, and three-dimensional model calculations.

Rinsland, Curtis P.↗

Analysis of a Generally Oriented Crack in a Functionally Graded Strip Sandwiched Between Two Homogeneous Half Planes

The driving forces for a generally oriented crack embedded in a Functionally Graded strip sandwiched between two half planes are analyzed using singular integral equations with Cauchy kernels, and integrated using Lobatto-Chebyshev collocation. Mixed-mode Stress Intensity Factors (SIF) and Strain Energy Release Rates (SERR) are calculated. The Stress Intensity Factors are compared for accuracy with previously published results. Parametric studies are conducted for various nonhomogeneity ratios, crack lengths. crack orientation and thickness of the strip. It is shown that the SERR is more complete and should be used for crack propagation analysis.

Shbeeb, N.↗

Timescales of Land Surface Evapotranspiration Response

Soil and vegetation exert strong control over the evapotranspiration rate, which couples the land surface water and energy balances. A method is presented to quantify the timescale of this surface control using daily general circulation model (GCM) simulation values of evapotranspiration and precipitation. By equating the time history of evaporation efficiency (ratio of actual to potential evapotranspiration) to the convolution of precipitation and a unit kernel (temporal weighting function), response functions are generated that can be used to characterize the timescales of evapotranspiration response for the land surface model (LSM) component of GCMS. The technique is applied to the output of two multiyear simulations of a GCM, one using a Surface-Vegetation-Atmosphere-Transfer (SVAT) scheme and the other a Bucket LSM. The derived response functions show that the Bucket LSM's response is significantly slower than that of the SVAT across the globe. The analysis also shows how the timescales of interception reservoir evaporation, bare soil evaporation, and vegetation transpiration differ within the SVAT LSM.

Scott, Russell↗

Testing New Programming Paradigms with NAS Parallel Benchmarks

Over the past decade, high performance computing has evolved rapidly, not only in hardware architectures but also with increasing complexity of real applications. Technologies have been developing to aim at scaling up to thousands of processors on both distributed and shared memory systems. Development of parallel programs on these computers is always a challenging task. Today, writing parallel programs with message passing (e.g. MPI) is the most popular way of achieving scalability and high performance. However, writing message passing programs is difficult and error prone. Recent years new effort has been made in defining new parallel programming paradigms. The best examples are: HPF (based on data parallelism) and OpenMP (based on shared memory parallelism). Both provide simple and clear extensions to sequential programs, thus greatly simplify the tedious tasks encountered in writing message passing programs. HPF is independent of memory hierarchy, however, due to the immaturity of compiler technology its performance is still questionable. Although use of parallel compiler directives is not new, OpenMP offers a portable solution in the shared-memory domain. Another important development involves the tremendous progress in the internet and its associated technology. Although still in its infancy, Java promisses portability in a heterogeneous environment and offers possibility to "compile once and run anywhere." In light of testing these new technologies, we implemented new parallel versions of the NAS Parallel Benchmarks (NPBs) with HPF and OpenMP directives, and extended the work with Java and Java-threads. The purpose of this study is to examine the effectiveness of alternative programming paradigms. NPBs consist of five kernels and three simulated applications that mimic the computation and data movement of large scale computational fluid dynamics (CFD) applications. We started with the serial version included in NPB2.3. Optimization of memory and cache usage was applied to several benchmarks, noticeably BT and SP, resulting in better sequential performance. In order to overcome the lack of an HPF performance model and guide the development of the HPF codes, we employed an empirical performance model for several primitives found in the benchmarks. We encountered a few limitations of HPF, such as lack of supporting the "REDISTRIBUTION" directive and no easy way to handle irregular computation. The parallelization with OpenMP directives was done at the outer-most loop level to achieve the largest granularity. The performance of six HPF and OpenMP benchmarks is compared with their MPI counterparts for the Class-A problem size in the figure in next page. These results were obtained on an SGI Origin2000 (195MHz) with MIPSpro-f77 compiler 7.2.1 for OpenMP and MPI codes and PGI pghpf-2.4.3 compiler with MPI interface for HPF programs.

Jin, H.↗

Ordering Unstructured Meshes for Sparse Matrix Computations on Leading Parallel Systems

The ability of computers to solve hitherto intractable problems and simulate complex processes using mathematical models makes them an indispensable part of modern science and engineering. Computer simulations of large-scale realistic applications usually require solving a set of non-linear partial differential equations (PDES) over a finite region. For example, one thrust area in the DOE Grand Challenge projects is to design future accelerators such as the SpaHation Neutron Source (SNS). Our colleagues at SLAC need to model complex RFQ cavities with large aspect ratios. Unstructured grids are currently used to resolve the small features in a large computational domain; dynamic mesh adaptation will be added in the future for additional efficiency. The PDEs for electromagnetics are discretized by the FEM method, which leads to a generalized eigenvalue problem Kx = AMx, where K and M are the stiffness and mass matrices, and are very sparse. In a typical cavity model, the number of degrees of freedom is about one million. For such large eigenproblems, direct solution techniques quickly reach the memory limits. Instead, the most widely-used methods are Krylov subspace methods, such as Lanczos or Jacobi-Davidson. In all the Krylov-based algorithms, sparse matrix-vector multiplication (SPMV) must be performed repeatedly. Therefore, the efficiency of SPMV usually determines the eigensolver speed. SPMV is also one of the most heavily used kernels in large-scale numerical simulations.

Oliker, Leonid↗

DMFS: A Data Migration File System for NetBSD

I have recently developed dmfs, a Data Migration File System, for NetBSD. This file system is based on the overlay file system, which is discussed in a separate paper, and provides kernel support for the data migration system being developed by my research group here at NASA/Ames. The file system utilizes an underlying file store to provide the file backing, and coordinates user and system access to the files. It stores its internal meta data in a flat file, which resides on a separate file system. Our data migration system provides archiving and file migration services. System utilities scan the dmfs file system for recently modified files, and archive them to two separate tape stores. Once a file has been doubly archived, files larger than a specified size will be truncated to that size, potentially freeing up large amounts of the underlying file store. Some sites will choose to retain none of the file (deleting its contents entirely from the file system) while others may choose to retain a portion, for instance a preamble describing the remainder of the file. The dmfs layer coordinates access to the file, retaining user-perceived access and modification times, file size, and restricting access to partially migrated files to the portion actually resident. When a user process attempts to read from the non-resident portion of a file, it is blocked and the dmfs layer sends a request to a system daemon to restore the file. As more of the file becomes resident, the user process is permitted to begin accessing the now-resident portions of the file. For simplicity, our data migration system divides a file into two portions, a resident portion followed by an optional non-resident portion. Also, a file is in one of three states: fully resident, fully resident and archived, and (partially) non-resident and archived. For a file which is only partially resident, any attempt to write or truncate the file, or to read a non-resident portion, will trigger a file restoration. Truncations and writes are blocked until the file is fully restored so that a restoration which only partially succeed does not leave the file in an indeterminate state with portions existing only on tape and other portions only in the disk file system. We chose layered file system technology as it permits us to focus on the data migration functionality, and permits end system administrators to choose the underlying file store technology. We chose the overlay layered file system instead of the null layer for two reasons: first to permit our layer to better preserve meta data integrity and second to prevent even root processes from accessing migrated files. This is achieved as the underlying file store becomes inaccessible once the dmfs layer is mounted. We are quite pleased with how the layered file system has turned out. Of the 45 vnode operations in NetBSD, 20 (forty-four percent) required no intervention by our file layer - they are passed directly to the underlying file store. Of the twenty five we do intercept, nine (such as vop_create()) are intercepted only to ensure meta data integrity. Most of the functionality was concentrated in five operations: vop_read, vop_write, vop_getattr, vop_setattr, and vop_fcntl. The first four are the core operations for controlling access to migrated files and preserving the user experience. vop_fcntl, a call generated for a certain class of fcntl codes, provides the command channel used by privileged user programs to communicate with the dmfs layer.

Studenmund, William↗

Aeroelastic Response of Nonlinear Wing Section by Functional Series Technique

This paper addresses the problem of the determination of the subcritical aeroelastic response and flutter instability of nonlinear two-dimensional lifting surfaces in an incompressible flow-field via indicial functions and Volterra series approach. The related aeroelastic governing equations are based upon the inclusion of structural and damping nonlinearities in plunging and pitching, of the linear unsteady aerodynamics and consideration of an arbitrary time-dependent external pressure pulse. Unsteady aeroelastic nonlinear kernels are determined, and based on these, frequency and time histories of the subcritical aeroelastic response are obtained, and in this context the influence of the considered nonlinearities is emphasized. Conclusions and results displaying the implications of the considered effects are supplied.

Silva, Walter A.↗

Volterra Series Approach for Nonlinear Aeroelastic Response of 2-D Lifting Surfaces

The problem of the determination of the subcritical aeroelastic response and flutter instability of nonlinear two-dimensional lifting surfaces in an incompressible flow-field via Volterra series approach is addressed. The related aeroelastic governing equations are based upon the inclusion of structural nonlinearities, of the linear unsteady aerodynamics and consideration of an arbitrary time-dependent external pressure pulse. Unsteady aeroelastic nonlinear kernels are determined, and based on these, frequency and time histories of the subcritical aeroelastic response are obtained, and in this context the influence of geometric nonlinearities is emphasized. Conclusions and results displaying the implications of the considered effects are supplied.

Silva, Walter A.↗

Aeroelastic Response of Nonlinear Wing Section By Functional Series Technique

This paper addresses the problem of the determination of the subcritical aeroelastic response and flutter instability of nonlinear two-dimensional lifting surfaces in an incompressible flow-field via indicial functions and Volterra series approach. The related aeroelastic governing equations are based upon the inclusion of structural and damping nonlinearities in plunging and pitching, of the linear unsteady aerodynamics and consideration of an arbitrary time-dependent external pressure pulse. Unsteady aeroelastic nonlinear kernels are determined, and based on these, frequency and time histories of the subcritical aeroelastic response are obtained, and in this context the influence of the considered nonlinearities is emphasized. Conclusions and results displaying the implications of the considered effects are supplied.

Marzocca, Piergiovanni↗

Applications Performance on NAS Intel Paragon XP/S - 15#

The Numerical Aerodynamic Simulation (NAS) Systems Division received an Intel Touchstone Sigma prototype model Paragon XP/S- 15 in February, 1993. The i860 XP microprocessor with an integrated floating point unit and operating in dual -instruction mode gives peak performance of 75 million floating point operations (NIFLOPS) per second for 64 bit floating point arithmetic. It is used in the Paragon XP/S-15 which has been installed at NAS, NASA Ames Research Center. The NAS Paragon has 208 nodes and its peak performance is 15.6 GFLOPS. Here, we will report on early experience using the Paragon XP/S- 15. We have tested its performance using both kernels and applications of interest to NAS. We have measured the performance of BLAS 1, 2 and 3 both assembly-coded and Fortran coded on NAS Paragon XP/S- 15. Furthermore, we have investigated the performance of a single node one-dimensional FFT, a distributed two-dimensional FFT and a distributed three-dimensional FFT Finally, we measured the performance of NAS Parallel Benchmarks (NPB) on the Paragon and compare it with the performance obtained on other highly parallel machines, such as CM-5, CRAY T3D, IBM SP I, etc. In particular, we investigated the following issues, which can strongly affect the performance of the Paragon: a. Impact of the operating system: Intel currently uses as a default an operating system OSF/1 AD from the Open Software Foundation. The paging of Open Software Foundation (OSF) server at 22 MB to make more memory available for the application degrades the performance. We found that when the limit of 26 NIB per node out of 32 MB available is reached, the application is paged out of main memory using virtual memory. When the application starts paging, the performance is considerably reduced. We found that dynamic memory allocation can help applications performance under certain circumstances. b. Impact of data cache on the i860/XP: We measured the performance of the BLAS both assembly coded and Fortran coded. We found that the measured performance of assembly-coded BLAS is much less than what memory bandwidth limitation would predict. The influence of data cache on different sizes of vectors is also investigated using one-dimensional FFTs. c. Impact of processor layout: There are several different ways processors can be laid out within the two-dimensional grid of processors on the Paragon. We have used the FFT example to investigate performance differences based on processors layout.

Saini, Subhash↗