Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75

Unsteady turbomachinery flow simulations on massively parallel architectures

The accurate numerical simulation of unsteady, three-dimensional viscous flow in turbomachines is computationally very intensive, requiring prohibitively large amounts of computer time on current vector supercomputers. In recent years, computer systems based on massively parallel architectures have been developed that offer the promise of meeting the computational power requirements of such large-scale simulations. However, a rethinking of existing algorithms and methodology is required in order to fully harness the computational power of such architectures. In this paper the capabilities of the Connection Machine (CM-2) in predicting unsteady flows in turbomachines are evaluated. The implementation on the CM-2 of an implicit, time-accurate, zonal algorithm for the Navier-Stokes equations in two dimensions is described. Programming issues and modifications made to the original algorithm (developed for vector, pipelined supercomputers) in order to improve performance on the CM-2 are outlined. Algorithm performance is evaluated and compared with a functionally equivalent code for the CRAY-YMP.

Madavan, N. K.↗

A Two Colorable Fourth Order Compact Difference Scheme and Parallel Iterative Solution of the 3D Convection Diffusion Equation

A new fourth order compact difference scheme for the three dimensional convection diffusion equation with variable coefficients is presented. The novelty of this new difference scheme is that it Only requires 15 grid points and that it can be decoupled with two colors. The entire computational grid can be updated in two parallel subsweeps with the Gauss-Seidel type iterative method. This is compared with the known 19 point fourth order compact differenCe scheme which requires four colors to decouple the computational grid. Numerical results, with multigrid methods implemented on a shared memory parallel computer, are presented to compare the 15 point and the 19 point fourth order compact schemes.

Zhang, Jun↗

Adaptive independent joint control of manipulators - Theory and experiment

The author presents a simple decentralized adaptive control scheme for multijoint robot manipulators based on the independent joint control concept. The proposed control scheme for each joint consists of a PID (proportional integral and differential) feedback controller and a position-velocity-acceleration feedforward controller, both with adjustable gains. The static and dynamic couplings that exist between the joint motions are compensated by the adaptive independent joint controllers while ensuring trajectory tracking. The proposed scheme is implemented on a MicroVAX II computer for motion control of the first three joints of a PUMA 560 arm. Experimental results are presented to demonstrate that trajectory tracking is achieved despite strongly coupled, highly nonlinear joint dynamics. The results confirm that the proposed decentralized adaptive control of manipulators is feasible, in spite of strong interactions between joint motions. The control scheme presented is computationally very fast and is amenable to parallel processing implementation within a distributed computing architecture, where each joint is controlled independently by a simple algorithm on a dedicated microprocessor.

Seraji, H.↗

ANALYSIS OF THE MSL/MEDLI ENTRY DATA WITH COUPLED CFD AND MATERIAL RESPONSE.

The Mars Science Laboratory (MSL) was protected during its atmospheric entry by an instrumented heat-shield using NASA's Phenolic Impregnated Carbon Ablator (PICA) material. PICA is a lightweight carbon fiber/polymeric resin material that offers out-standing performances for protecting probes during planetary entry. The Mars Entry Descent and Landing Instrument (MEDLI) suite on MSL offers unique in-flight validation data for models of material response and atmospheric entry. MEDLI recorded, among other things, time-resolved in-depth temperature data of PICA using thermocouple sensors assembled in the MEDLI Integrated Sensor Plugs (MISP). The objective of this work is to showcase and analyze the coupling between the material response and the aerothermal environment. As shown in Figure 1, the workflow is divided into the following steps. First, the aerothermal properties are computed in the Data Parallel Line Relaxation (DPLR) code [3] and used with the Nonequilibrium air radiation (NEQAIR) program [8] to compute radiative heating. Second, the thermal response inside the material is computed in the Porous material Analysis Toolbox based on Open-FOAM (PATO) using a fixed blowing correction parameter. Third, the pyrolysis gases computed in PATO are used as inputs to a blowing boundary condition within DPLR. Fourth, the new environment properties from DPLR are used in NEQAIR to provide an updated solution, then both the updated aerothermal environment and radiative heating are used in PATO without blowing correction. The third and fourth steps are then repeated until convergence in surface temperature is obtained. Convergence in the radiative heating is generally achieved before surface temperature, at which point the radiative heating is no longer updated. Char mass loss rates are forced to zero to produce a non-receding surface condition. For early time points in the trajectory, where flow around the MSL aeroshell is rarefied, the Direct Simulation Monte Carlo (DSMC) code, SPARTA, is used to compute the aerothermal environment. Iteration between PATO and SPARTA is not performed due to the computational cost of DSMC simulations. Preliminary results of the coupling between PATO and DPLR for the MSL heatshield atmospheric entry model are presented in Figures 2-4 at 65 seconds after entry interface. Figure 2 shows the surface temperature results from an uncoupled simulation in PATO with the blowing correction parameter applied (left) along with the coupled surface temperature after iteration (right). Figure 3 shows the surface temperature along the centerline from windward to leeward for easier comparison. Figure 4 shows the coupled and uncoupled pyrolysis gas blowing rate. Mars 2020 used a similar heatshield consisting of PICA for thermal protection during entry, descent, and landing. In preparation for Mars 2020 post-flight analysis, the predictive material response capability is benchmarked against flight data from MEDLI. This work represents an important milestone toward the development of validated predictive capabilities for designing thermal protection systems for planetary probes.

Mars Science Laboratory↗

Analysis of MSL/MEDLI Entry Data with Coupled CFD and Material Response

The Mars Science Laboratory (MSL) was protected during its atmospheric entry by an instrumented heatshield using NASA's Phenolic Impregnated Carbon Ablator (PICA) material [1]. PICA is a lightweight carbon fiber/polymeric resin material that offers outstanding performances for protecting probes during planetary entry. The Mars Entry Descent and Landing Instrument (MEDLI) suite on MSL offers unique in-flight validation data for models of material response and atmospheric entry. MEDLI recorded, among other things, time-resolved in-depth temperature data of PICA using thermocouple sensors assembled in the MEDLI Integrated Sensor Plugs (MISP) [2]. The objective of this work is to showcase and analyze the coupling between the material response and the aerothermal environment. As shown in Figure 1, the workflow is divided into the following steps. First, the aerothermal properties are computed in the Data Parallel Line Relaxation (DPLR) code [3] and used with the Nonequilibrium air radiation (NEQAIR) program [8] to compute radiative heating. Second, the thermal response inside the material is computed in the Porous material Analysis Toolbox based on OpenFOAM (PATO) [4,5,6] using a fixed blowing correction parameter. Third, the pyrolysis gases computed in PATO are used as inputs to a blowing boundary condition within DPLR. Fourth, the new environment properties from DPLR are used in NEQAIR to provide an updated solution, then both the updated aerothermal environment and radiative heating are used in PATO without blowing correction. The third and fourth steps are then repeated until convergence in surface temperature is obtained. Convergence in the radiative heating is generally achieved before surface temperature, at which point the radiative heating is no longer updated. Char mass loss rates are forced to zero to produce a non-receding surface condition. For early time points in the trajectory, where flow around the MSL aeroshell is rarefied, the Direct Simulation Monte Carlo (DSMC) code, SPARTA [7], is used to compute the aerothermal environment. Iteration between PATO and SPARTA is not performed due to the computational cost of DSMC simulations. Preliminary results of the coupling between PATO and DPLR for the MSL heatshield atmospheric entry model are presented in Figures 2-4 at 65 seconds after entry interface. Figure 2 shows the surface temperature results from an uncoupled simulation in PATO with the blowing correction parameter applied (left) along with the coupled surface temperature after iteration (right). Figure 3 shows the surface temperature along the centerline from windward to leeward for easier comparison. Figure 4 shows the coupled and uncoupled pyrolysis gas blowing rate. Mars 2020 used a similar heatshield consisting of PICA for thermal protection during entry, descent, and landing. In preparation for Mars 2020 post-flight analysis, the predictive material response capability is benchmarked against flight data from MEDLI. This work represents an important milestone toward the development of validated predictive capabilities for designing thermal protection systems for planetary probes.

Thermal Protection Systems↗

Computer program MCAP-TOSS calculates steady-state fluid dynamics of coolant in parallel channels and temperature distribution in surrounding heat-generating solid

Computer program calculates the steady state fluid distribution, temperature rise, and pressure drop of a coolant, the material temperature distribution of a heat generating solid, and the heat flux distributions at the fluid-solid interfaces. It performs the necessary iterations automatically within the computer, in one machine run.

Lee, A. Y.↗

Probabilistic Design of a Wind Tunnel Model to Match the Response of a Full-Scale Aircraft

approach is presented for carrying out the reliability-based design of a plate-like wing that is part of a wind tunnel model. The goal is to design the wind tunnel model to match the stiffness characteristics of the wing box of a flight vehicle while satisfying strength-based risk/reliability requirements that prevents damage to the wind tunnel model and fixtures. The flight vehicle is a modified F/A-18 aircraft. The design problem is solved using reliability-based optimization techniques. The objective function to be minimized is the difference between the displacements of the wind tunnel model and the corresponding displacements of the flight vehicle. The design variables control the thickness distribution of the wind tunnel model. Displacements of the wind tunnel model change with the thickness distribution, while displacements of the flight vehicle are a set of fixed data. The only constraint imposed is that the probability of failure is less than a specified value. Failure is assumed to occur if the stress caused by aerodynamic pressure loading is greater than the specified strength allowable. Two uncertain quantities are considered: the allowable stress and the thickness distribution of the wind tunnel model. Reliability is calculated using Monte Carlo simulation with response surfaces that provide approximate values of stresses. The response surface equations are, in turn, computed from finite element analyses of the wind tunnel model at specified design points. Because the response surface approximations were fit over a small region centered about the current design, the response surfaces were refit periodically as the design variables changed. Coarse-grained parallelism was used to simultaneously perform multiple finite element analyses. Studies carried out in this paper demonstrate that this scheme of using moving response surfaces and coarse-grained computational parallelism reduce the execution time of the Monte Carlo simulation enough to make the design problem tractable. The results of the reliability-based designs performed in this paper show that large decreases in the probability of stress-based failure can be realized with only small sacrifices in the ability of the wind tunnel model to represent the displacements of the full-scale vehicle.

Mason, Brian H.↗

Optimal processor assignment for pipeline computations

The availability of large scale multitasked parallel architectures introduces the following processor assignment problem for pipelined computations. Given a set of tasks and their precedence constraints, along with their experimentally determined individual responses times for different processor sizes, find an assignment of processor to tasks. Two objectives are of interest: minimal response given a throughput requirement, and maximal throughput given a response time requirement. These assignment problems differ considerably from the classical mapping problem in which several tasks share a processor; instead, it is assumed that a large number of processors are to be assigned to a relatively small number of tasks. Efficient assignment algorithms were developed for different classes of task structures. For a p processor system and a series parallel precedence graph with n constituent tasks, an O(np2) algorithm is provided that finds the optimal assignment for the response time optimization problem; it was found that the assignment optimizing the constrained throughput in O(np2log p) time. Special cases of linear, independent, and tree graphs are also considered.

Nicol, David M.↗

A Partitioned - Task Parallel Implementation of the NASA Multiscale Analysis Tool for High Performance Computing

The NASA Multiscale Analysis Tool (NASMAT) is a platform for multiscale modeling of composites which can perform analysis of materials with any arbitrary number of length scales. The platform supports modularity, scalability, and interoperability using recursive procedures and data structures. A Macro solver driven parallelization scheme often limits the capability of NASMAT to scale as it has access to limited memory and number of cores (often one core/thread) and often forces to implement macro solver specific changes to the platform. In this work, a partitioned task-parallel approach is adopted, where the parallelization strategy adopted for NASMAT is independent of the macro solver and the computational resources are managed independently. The programming architecture takes into account the hierarchy of multiple scales (task-dependence) and the heterogeneous nature (dynamic load balancing) of computation through implementation of a hierarchy-informed task parallel model. The partitioned nature of the framework further extends the “plug and play” capability of NASMAT. preCICE, an open-source library for coupling multiphysics solver in a partitioned manner, is adopted to integrate NASMAT with an external macro solver by implementing a NASMAT adapter for preCICE. Speedup and scalability of the framework is studied for micromechanical models of varying size.

task-parallel↗

Parallel-vector out-of-core equation solver for computational mechanics

A parallel/vector out-of-core equation solver is developed for shared-memory computers, such as the Cray Y-MP machine. The input/ output (I/O) time is reduced by using the a synchronous BUFFER IN and BUFFER OUT, which can be executed simultaneously with the CPU instructions. The parallel and vector capability provided by the supercomputers is also exploited to enhance the performance. Numerical applications in large-scale structural analysis are given to demonstrate the efficiency of the present out-of-core solver.

Qin, J.↗

Head-on parallel blade-vortex interaction

An experimental and computational study was carried out to investigate the parallel head-on blade-vortex interaction (BVI) and its noise generation mechanism. A shock tube, with an enlarged test section, was used to generate a compressible starting vortex which interacted with a target airfoil. The dual-pulsed holographic interferometry (DPHI) technique and airfoil surface pressure measurements were employed to obtain quantitative flow data during the BVI. A thin-layer Navier-Stokes code (BV12D), with a high-order upwind-biased scheme and a multizonal grid, was also used to simulate numerically the phenomena occurring in the head-on BVI. The detailed structure of a convecting vortex was studied through independent measurements of density and pressure distributions across the vortex center. Results indicate that, in a strong head-on BVI, the opposite pressure peaks are generated on both sides of the leading edge as the vortex approaches. Then, as soon as the vortex passes by the leading edge, the high-pressure peak suddenly moves toward the low-peak-reducing in magnitude as it moves--simultaneously giving rise to the initial sound wave. In both experiment and computation, it is shown that the viscous effect plays a significant role in head-on BVIs.

Lee, Soogab↗

Scalability of GlennICE in a Parallel Environment

GlennICE (Glenn Icing Computational Environment) is a comptational tool designed to calculate ice growth on complex three- dimensional geometries using the input from a user-supplied computational fluid dynamics (CFD) solution for the geometry of interest. The most significant developments in the advancement of GlennICE have been investigating the convergence of the collection efficiency, efficiently finding trajectories, and improving the refinement methodology. Such developments have increased the efficiency of GlennICE for tractability in a practical engineering application. Although studies have demonstrated a reduction in the amount of work (memory footprint) required, research has yet to systematically investigate the effects of scaling GlennICE. This paper sets out to benchmark the scalability of GlennICE within a parallel environment and investigate if an increase in the number of processors result in a linear speed up.

Computational Icing↗

Accessing and Visualizing scientific spatiotemporal data

This paper discusses work done by JPL 's Parallel Applications Technologies Group in helping scientists access and visualize very large data sets through the use of multiple computing resources, such as parallel supercomputers, clusters, and grids These tools do one or more of the following tasks visualize local data sets for local users, visualize local data sets for remote users, and access and visualize remote data sets The tools are used for various types of data, including remotely sensed image data, digital elevation models, astronomical surveys, etc The paper attempts to pull some common elements out of these tools that may be useful for others who have to work with similarly large data sets.

data sets↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Application of data flow concepts to a multigrid solver for the Euler equations

In this study a multigrid solver for Euler equations (FLO52R) was examined to determine its performance potential on a hypothetical computer using a data flow architecture. The proposed computer would require massive parallelism to realize its design performance. On the other hand this parallelism would be more easily realized than with a conventional vector processor such as the Cray-1S. Several changes to the proposed design substantially alleviated most of the remaining bottlenecks to parallel processing. Other changes allowed clearer definition of memory access and disk I/O. Finally, a portion of the algorithm was rewritten to improve parallel performance. With these changes, performance levels approaching that of a Cray-1S may be possible for a computer costing far less. Estimates are given for overall speed, memory, and network bandwidth, and for instruction memory requirements.

Merriam, M. L.↗

A distributed Clips implementation: dClips

A distributed version of the Clips language, dClips, was implemented on top of two existing generic distributed messaging systems to show that: (1) it is easy to create a coarse-grained parallel programming environment out of an existing language if a high level messaging system is used; and (2) the computing model of a parallel programming environment can be changed easily if we change the underlying messaging system. dClips processes were first connected with a simple master-slave model. A client-server model with intercommunicating agents was later implemented. The concept of service broker is being investigated.

Li, Y. Philip↗

Accessing and visualizing scientific spatiotemporal data

This paper discusses work done by JPL's Parallel Applications Technologies Group in helping scientists access and visualize very large data sets through the use of multiple computing resources, such as parallel supercomputers, clusters, and grids.

rendering↗