Search NASA⌕ Search

SEARCH · Search NASA

Results for “parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69

Implementation of a Parallel Kalman Filter for Stratospheric Chemical Tracer Assimilation

A Kalman filter for the assimilation of long-lived atmospheric chemical constituents has been developed for two-dimensional transport models on isentropic surfaces over the globe. An important attribute of the Kalman filter is that it calculates error covariances of the constituent fields using the tracer dynamics. Consequently, the current Kalman-filter assimilation is a five-dimensional problem (coordinates of two points and time), and it can only be handled on computers with large memory and high floating point speed. In this paper, an implementation of the Kalman filter for distributed-memory, message-passing parallel computers is discussed. Two approaches were studied: an operator decomposition and a covariance decomposition. The latter was found to be more scalable than the former, and it possesses the property that the dynamical model does not need to be parallelized, which is of considerable practical advantage. This code is currently used to assimilate constituent data retrieved by limb sounders on the Upper Atmosphere Research Satellite. Tests of the code examined the variance transport and observability properties. Aspects of the parallel implementation, some timing results, and a brief discussion of the physical results will be presented.

Chang, Lang-Ping↗

Predicting Flows of Rarefied Gases

DSMC Analysis Code (DAC) is a flexible, highly automated, easy-to-use computer program for predicting flows of rarefied gases -- especially flows of upper-atmospheric, propulsion, and vented gases impinging on spacecraft surfaces. DAC implements the direct simulation Monte Carlo (DSMC) method, which is widely recognized as standard for simulating flows at densities so low that the continuum-based equations of computational fluid dynamics are invalid. DAC enables users to model complex surface shapes and boundary conditions quickly and easily. The discretization of a flow field into computational grids is automated, thereby relieving the user of a traditionally time-consuming task while ensuring (1) appropriate refinement of grids throughout the computational domain, (2) determination of optimal settings for temporal discretization and other simulation parameters, and (3) satisfaction of the fundamental constraints of the method. In so doing, DAC ensures an accurate and efficient simulation. In addition, DAC can utilize parallel processing to reduce computation time. The domain decomposition needed for parallel processing is completely automated, and the software employs a dynamic load-balancing mechanism to ensure optimal parallel efficiency throughout the simulation.

LeBeau, Gerald J.↗

Lanczos eigensolution method for high-performance computers

The theory, computational analysis, and applications are presented of a Lanczos algorithm on high performance computers. The computationally intensive steps of the algorithm are identified as: the matrix factorization, the forward/backward equation solution, and the matrix vector multiples. These computational steps are optimized to exploit the vector and parallel capabilities of high performance computers. The savings in computational time from applying optimization techniques such as: variable band and sparse data storage and access, loop unrolling, use of local memory, and compiler directives are presented. Two large scale structural analysis applications are described: the buckling of a composite blade stiffened panel with a cutout, and the vibration analysis of a high speed civil transport. The sequential computational time for the panel problem executed on a CONVEX computer of 181.6 seconds was decreased to 14.1 seconds with the optimized vector algorithm. The best computational time of 23 seconds for the transport problem with 17,000 degs of freedom was on the the Cray-YMP using an average of 3.63 processors.

Bostic, Susan W.↗

Unsteady turbomachinery flow simulations on massively parallel architectures

The accurate numerical simulation of unsteady, three-dimensional viscous flow in turbomachines is computationally very intensive, requiring prohibitively large amounts of computer time on current vector supercomputers. In recent years, computer systems based on massively parallel architectures have been developed that offer the promise of meeting the computational power requirements of such large-scale simulations. However, a rethinking of existing algorithms and methodology is required in order to fully harness the computational power of such architectures. In this paper the capabilities of the Connection Machine (CM-2) in predicting unsteady flows in turbomachines are evaluated. The implementation on the CM-2 of an implicit, time-accurate, zonal algorithm for the Navier-Stokes equations in two dimensions is described. Programming issues and modifications made to the original algorithm (developed for vector, pipelined supercomputers) in order to improve performance on the CM-2 are outlined. Algorithm performance is evaluated and compared with a functionally equivalent code for the CRAY-YMP.

Madavan, N. K.↗

A Two Colorable Fourth Order Compact Difference Scheme and Parallel Iterative Solution of the 3D Convection Diffusion Equation

A new fourth order compact difference scheme for the three dimensional convection diffusion equation with variable coefficients is presented. The novelty of this new difference scheme is that it Only requires 15 grid points and that it can be decoupled with two colors. The entire computational grid can be updated in two parallel subsweeps with the Gauss-Seidel type iterative method. This is compared with the known 19 point fourth order compact differenCe scheme which requires four colors to decouple the computational grid. Numerical results, with multigrid methods implemented on a shared memory parallel computer, are presented to compare the 15 point and the 19 point fourth order compact schemes.

Zhang, Jun↗

Adaptive independent joint control of manipulators - Theory and experiment

The author presents a simple decentralized adaptive control scheme for multijoint robot manipulators based on the independent joint control concept. The proposed control scheme for each joint consists of a PID (proportional integral and differential) feedback controller and a position-velocity-acceleration feedforward controller, both with adjustable gains. The static and dynamic couplings that exist between the joint motions are compensated by the adaptive independent joint controllers while ensuring trajectory tracking. The proposed scheme is implemented on a MicroVAX II computer for motion control of the first three joints of a PUMA 560 arm. Experimental results are presented to demonstrate that trajectory tracking is achieved despite strongly coupled, highly nonlinear joint dynamics. The results confirm that the proposed decentralized adaptive control of manipulators is feasible, in spite of strong interactions between joint motions. The control scheme presented is computationally very fast and is amenable to parallel processing implementation within a distributed computing architecture, where each joint is controlled independently by a simple algorithm on a dedicated microprocessor.

Seraji, H.↗

ANALYSIS OF THE MSL/MEDLI ENTRY DATA WITH COUPLED CFD AND MATERIAL RESPONSE.

The Mars Science Laboratory (MSL) was protected during its atmospheric entry by an instrumented heat-shield using NASA's Phenolic Impregnated Carbon Ablator (PICA) material. PICA is a lightweight carbon fiber/polymeric resin material that offers out-standing performances for protecting probes during planetary entry. The Mars Entry Descent and Landing Instrument (MEDLI) suite on MSL offers unique in-flight validation data for models of material response and atmospheric entry. MEDLI recorded, among other things, time-resolved in-depth temperature data of PICA using thermocouple sensors assembled in the MEDLI Integrated Sensor Plugs (MISP). The objective of this work is to showcase and analyze the coupling between the material response and the aerothermal environment. As shown in Figure 1, the workflow is divided into the following steps. First, the aerothermal properties are computed in the Data Parallel Line Relaxation (DPLR) code [3] and used with the Nonequilibrium air radiation (NEQAIR) program [8] to compute radiative heating. Second, the thermal response inside the material is computed in the Porous material Analysis Toolbox based on Open-FOAM (PATO) using a fixed blowing correction parameter. Third, the pyrolysis gases computed in PATO are used as inputs to a blowing boundary condition within DPLR. Fourth, the new environment properties from DPLR are used in NEQAIR to provide an updated solution, then both the updated aerothermal environment and radiative heating are used in PATO without blowing correction. The third and fourth steps are then repeated until convergence in surface temperature is obtained. Convergence in the radiative heating is generally achieved before surface temperature, at which point the radiative heating is no longer updated. Char mass loss rates are forced to zero to produce a non-receding surface condition. For early time points in the trajectory, where flow around the MSL aeroshell is rarefied, the Direct Simulation Monte Carlo (DSMC) code, SPARTA, is used to compute the aerothermal environment. Iteration between PATO and SPARTA is not performed due to the computational cost of DSMC simulations. Preliminary results of the coupling between PATO and DPLR for the MSL heatshield atmospheric entry model are presented in Figures 2-4 at 65 seconds after entry interface. Figure 2 shows the surface temperature results from an uncoupled simulation in PATO with the blowing correction parameter applied (left) along with the coupled surface temperature after iteration (right). Figure 3 shows the surface temperature along the centerline from windward to leeward for easier comparison. Figure 4 shows the coupled and uncoupled pyrolysis gas blowing rate. Mars 2020 used a similar heatshield consisting of PICA for thermal protection during entry, descent, and landing. In preparation for Mars 2020 post-flight analysis, the predictive material response capability is benchmarked against flight data from MEDLI. This work represents an important milestone toward the development of validated predictive capabilities for designing thermal protection systems for planetary probes.

Mars Science Laboratory↗

Analysis of MSL/MEDLI Entry Data with Coupled CFD and Material Response

The Mars Science Laboratory (MSL) was protected during its atmospheric entry by an instrumented heatshield using NASA's Phenolic Impregnated Carbon Ablator (PICA) material [1]. PICA is a lightweight carbon fiber/polymeric resin material that offers outstanding performances for protecting probes during planetary entry. The Mars Entry Descent and Landing Instrument (MEDLI) suite on MSL offers unique in-flight validation data for models of material response and atmospheric entry. MEDLI recorded, among other things, time-resolved in-depth temperature data of PICA using thermocouple sensors assembled in the MEDLI Integrated Sensor Plugs (MISP) [2]. The objective of this work is to showcase and analyze the coupling between the material response and the aerothermal environment. As shown in Figure 1, the workflow is divided into the following steps. First, the aerothermal properties are computed in the Data Parallel Line Relaxation (DPLR) code [3] and used with the Nonequilibrium air radiation (NEQAIR) program [8] to compute radiative heating. Second, the thermal response inside the material is computed in the Porous material Analysis Toolbox based on OpenFOAM (PATO) [4,5,6] using a fixed blowing correction parameter. Third, the pyrolysis gases computed in PATO are used as inputs to a blowing boundary condition within DPLR. Fourth, the new environment properties from DPLR are used in NEQAIR to provide an updated solution, then both the updated aerothermal environment and radiative heating are used in PATO without blowing correction. The third and fourth steps are then repeated until convergence in surface temperature is obtained. Convergence in the radiative heating is generally achieved before surface temperature, at which point the radiative heating is no longer updated. Char mass loss rates are forced to zero to produce a non-receding surface condition. For early time points in the trajectory, where flow around the MSL aeroshell is rarefied, the Direct Simulation Monte Carlo (DSMC) code, SPARTA [7], is used to compute the aerothermal environment. Iteration between PATO and SPARTA is not performed due to the computational cost of DSMC simulations. Preliminary results of the coupling between PATO and DPLR for the MSL heatshield atmospheric entry model are presented in Figures 2-4 at 65 seconds after entry interface. Figure 2 shows the surface temperature results from an uncoupled simulation in PATO with the blowing correction parameter applied (left) along with the coupled surface temperature after iteration (right). Figure 3 shows the surface temperature along the centerline from windward to leeward for easier comparison. Figure 4 shows the coupled and uncoupled pyrolysis gas blowing rate. Mars 2020 used a similar heatshield consisting of PICA for thermal protection during entry, descent, and landing. In preparation for Mars 2020 post-flight analysis, the predictive material response capability is benchmarked against flight data from MEDLI. This work represents an important milestone toward the development of validated predictive capabilities for designing thermal protection systems for planetary probes.

Thermal Protection Systems↗

Computer program MCAP-TOSS calculates steady-state fluid dynamics of coolant in parallel channels and temperature distribution in surrounding heat-generating solid

Computer program calculates the steady state fluid distribution, temperature rise, and pressure drop of a coolant, the material temperature distribution of a heat generating solid, and the heat flux distributions at the fluid-solid interfaces. It performs the necessary iterations automatically within the computer, in one machine run.

Lee, A. Y.↗

Probabilistic Design of a Wind Tunnel Model to Match the Response of a Full-Scale Aircraft

approach is presented for carrying out the reliability-based design of a plate-like wing that is part of a wind tunnel model. The goal is to design the wind tunnel model to match the stiffness characteristics of the wing box of a flight vehicle while satisfying strength-based risk/reliability requirements that prevents damage to the wind tunnel model and fixtures. The flight vehicle is a modified F/A-18 aircraft. The design problem is solved using reliability-based optimization techniques. The objective function to be minimized is the difference between the displacements of the wind tunnel model and the corresponding displacements of the flight vehicle. The design variables control the thickness distribution of the wind tunnel model. Displacements of the wind tunnel model change with the thickness distribution, while displacements of the flight vehicle are a set of fixed data. The only constraint imposed is that the probability of failure is less than a specified value. Failure is assumed to occur if the stress caused by aerodynamic pressure loading is greater than the specified strength allowable. Two uncertain quantities are considered: the allowable stress and the thickness distribution of the wind tunnel model. Reliability is calculated using Monte Carlo simulation with response surfaces that provide approximate values of stresses. The response surface equations are, in turn, computed from finite element analyses of the wind tunnel model at specified design points. Because the response surface approximations were fit over a small region centered about the current design, the response surfaces were refit periodically as the design variables changed. Coarse-grained parallelism was used to simultaneously perform multiple finite element analyses. Studies carried out in this paper demonstrate that this scheme of using moving response surfaces and coarse-grained computational parallelism reduce the execution time of the Monte Carlo simulation enough to make the design problem tractable. The results of the reliability-based designs performed in this paper show that large decreases in the probability of stress-based failure can be realized with only small sacrifices in the ability of the wind tunnel model to represent the displacements of the full-scale vehicle.

Mason, Brian H.↗

Optimal processor assignment for pipeline computations

The availability of large scale multitasked parallel architectures introduces the following processor assignment problem for pipelined computations. Given a set of tasks and their precedence constraints, along with their experimentally determined individual responses times for different processor sizes, find an assignment of processor to tasks. Two objectives are of interest: minimal response given a throughput requirement, and maximal throughput given a response time requirement. These assignment problems differ considerably from the classical mapping problem in which several tasks share a processor; instead, it is assumed that a large number of processors are to be assigned to a relatively small number of tasks. Efficient assignment algorithms were developed for different classes of task structures. For a p processor system and a series parallel precedence graph with n constituent tasks, an O(np2) algorithm is provided that finds the optimal assignment for the response time optimization problem; it was found that the assignment optimizing the constrained throughput in O(np2log p) time. Special cases of linear, independent, and tree graphs are also considered.

Nicol, David M.↗

A Partitioned - Task Parallel Implementation of the NASA Multiscale Analysis Tool for High Performance Computing

The NASA Multiscale Analysis Tool (NASMAT) is a platform for multiscale modeling of composites which can perform analysis of materials with any arbitrary number of length scales. The platform supports modularity, scalability, and interoperability using recursive procedures and data structures. A Macro solver driven parallelization scheme often limits the capability of NASMAT to scale as it has access to limited memory and number of cores (often one core/thread) and often forces to implement macro solver specific changes to the platform. In this work, a partitioned task-parallel approach is adopted, where the parallelization strategy adopted for NASMAT is independent of the macro solver and the computational resources are managed independently. The programming architecture takes into account the hierarchy of multiple scales (task-dependence) and the heterogeneous nature (dynamic load balancing) of computation through implementation of a hierarchy-informed task parallel model. The partitioned nature of the framework further extends the “plug and play” capability of NASMAT. preCICE, an open-source library for coupling multiphysics solver in a partitioned manner, is adopted to integrate NASMAT with an external macro solver by implementing a NASMAT adapter for preCICE. Speedup and scalability of the framework is studied for micromechanical models of varying size.

task-parallel↗

Parallel-vector out-of-core equation solver for computational mechanics

A parallel/vector out-of-core equation solver is developed for shared-memory computers, such as the Cray Y-MP machine. The input/ output (I/O) time is reduced by using the a synchronous BUFFER IN and BUFFER OUT, which can be executed simultaneously with the CPU instructions. The parallel and vector capability provided by the supercomputers is also exploited to enhance the performance. Numerical applications in large-scale structural analysis are given to demonstrate the efficiency of the present out-of-core solver.

Qin, J.↗

Head-on parallel blade-vortex interaction

An experimental and computational study was carried out to investigate the parallel head-on blade-vortex interaction (BVI) and its noise generation mechanism. A shock tube, with an enlarged test section, was used to generate a compressible starting vortex which interacted with a target airfoil. The dual-pulsed holographic interferometry (DPHI) technique and airfoil surface pressure measurements were employed to obtain quantitative flow data during the BVI. A thin-layer Navier-Stokes code (BV12D), with a high-order upwind-biased scheme and a multizonal grid, was also used to simulate numerically the phenomena occurring in the head-on BVI. The detailed structure of a convecting vortex was studied through independent measurements of density and pressure distributions across the vortex center. Results indicate that, in a strong head-on BVI, the opposite pressure peaks are generated on both sides of the leading edge as the vortex approaches. Then, as soon as the vortex passes by the leading edge, the high-pressure peak suddenly moves toward the low-peak-reducing in magnitude as it moves--simultaneously giving rise to the initial sound wave. In both experiment and computation, it is shown that the viscous effect plays a significant role in head-on BVIs.

Lee, Soogab↗

Scalability of GlennICE in a Parallel Environment

GlennICE (Glenn Icing Computational Environment) is a comptational tool designed to calculate ice growth on complex three- dimensional geometries using the input from a user-supplied computational fluid dynamics (CFD) solution for the geometry of interest. The most significant developments in the advancement of GlennICE have been investigating the convergence of the collection efficiency, efficiently finding trajectories, and improving the refinement methodology. Such developments have increased the efficiency of GlennICE for tractability in a practical engineering application. Although studies have demonstrated a reduction in the amount of work (memory footprint) required, research has yet to systematically investigate the effects of scaling GlennICE. This paper sets out to benchmark the scalability of GlennICE within a parallel environment and investigate if an increase in the number of processors result in a linear speed up.

Computational Icing↗