Search NASA⌕ Search

SEARCH · Search NASA

Results for “Parallel Programming”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Introduction to: Atlantic Meridional Overturning Circulation(AMOC)

A striking conclusion of the Intergovernmental Panel on Climate Change 2007 report is the crucial role that the Atlantic Meridional Overturning Circulation (AMOC) may play in anthropogenic climate change. However, these IPCC coupled climate simulations show a broad range of uncertainty in the magnitude and timing of AMOC transport change ranging from none to nearly complete collapse within the 21st century. The potential consequences of large changes in the characteristics of AMOC have motivated the creation in the United States of an interagency program and implementation plan to develop monitoring and prediction capabilities for the AMOC This program parallels the development of substantial monitoring efforts by European, South American and African countries -- notably the UK Rapid and Rapid-Watch programs. The papers contained in this volume are derived from presentations at the First U.S. Atlantic Meridional Overturning Circulation (AMOC) Meeting held 4 - 6 May, 2009 to review the US implementation plan and its coordination with other monitoring activities. The Atlantic Meridional Overturning Circulation consists of multiple components illustrated in an attached figure. Water enters the South Atlantic at upper and intermediate depths through both western and eastern routes (where eddy transport is especially important) and is transported northward across the equator, where it recirculates within the northern subtropical and subpolar gyres. The northern end is defined by the sinking regions of the Nordic Seas and the Labrador Sea where the waters that eventually form the upper and lower branches of North Atlantic Deep Water are conditioned. High surface salinities, the result of high net evaporation in the tropics and subtropics (including the Mediterranean Sea), and presence of regions of the Arctic Ocean that remain ice-free even in winter allow for the rapid cooling and thus densification of surface water. This dense surface water becomes the source of deep water formation in the sinking regions. In addition to transporting mass, the AMOC transports roughly half of the total amount of heat carried northward through the northern subtropics (down the temperature-gradient) by the ocean. In contrast in the Southern Hemisphere AMOC transports heat up-gradient from the cool Circumpolar Current to the warm tropics. Paleoevidence suggests that AMOC heat transport in the two hemispheres has varied over time in ways intimately tied to millennial changes in the Earth's climate. In one example, the abrupt Younger Dryas spell of cold weather over the North Atlantic, which began 13,000 years ago, has generally been linked to a millennial shutdown of the AMOC as a result of massive freshwater discharge from the North American continent. The current AMOC monitoring array consists of a series of instrumented transects located across key passages (see Cunningham et al., 2010 for a recent review). In the Arctic and sub-Arctic, transects cross Fram Strait, Denmark Strait and the Faroe Channel (connecting Greenland, Iceland, and the United Kingdom), as well as the entrance to the Labrador Sea. Further south and extending outwards from the east coast of North America there are a series of monitoring arrays including arrays of the Canadian Atlantic Zone Monitoring Program, deployments of the Rapid Western Atlantic Variability Experiment (WAVE), Line W at 39 N, as well as the Rapid-MOC moored array. The latter spans the entire Atlantic basin along 26.5 N. At tropical latitudes we have the Meridional Overturning Variability Experiment (MOVE) array at 16 N, while in the Southern Hemisphere a corresponding basin-spanning transect is being established at the latitude of Cape of Good Hope, complemented by arrays at Drake Passage.

Hakkinen, Sirpa↗

MTF Driven by Plasma Liner Dynamically Formed by the Merging of Plasma Jets: An Overview

One approach for standoff delivery of the momentum flux for compressing the target in MTF consists of using a spherical array of plasma jets to form a spherical plasma shell imploding towards the center of a magnetized plasma, a compact toroid (Figure 1). A 3-year experiment (PLX-1) to explore the physics of forming a 2-D plasma liner (shell) by merging plasma jets is described. An overview showing how this 3-year project (PLX-1) fits into the program plan at the national and international level for realizing MTF for energy and propulsion is discussed. Assuming that there will be a parallel program in demonstrating and establishing the underlying physics principles of MTF using whatever liner is appropriate (e.g. a solid liner) with a goal of demonstrating breakeven by 2010, the current research effort at NASA MSFC attempts to complement such a program by addressing the issues of practical embodiment of MTF for propulsion. Successful conclusion of PLX-1 will be followed by a Physics Feasibility Experiment (PLX-2) for the Plasma Liner Driven MTF.

Thio, Y. C. Francis↗

Message Passing and Shared Address Space Parallelism on an SMP Cluster

Currently, message passing (MP) and shared address space (SAS) are the two leading parallel programming paradigms. MP has been standardized with MPI, and is the more common and mature approach; however, code development can be extremely difficult, especially for irregularly structured computations. SAS offers substantial ease of programming, but may suffer from performance limitations due to poor spatial locality and high protocol overhead. In this paper, we compare the performance of and the programming effort required for six applications under both programming models on a 32-processor PC-SMP cluster, a platform that is becoming increasingly attractive for high-end scientific computing. Our application suite consists of codes that typically do not exhibit scalable performance under shared-memory programming due to their high communication-to-computation ratios and/or complex communication patterns. Results indicate that SAS can achieve about half the parallel efficiency of MPI for most of our applications, while being competitive for the others. A hybrid MPI+SAS strategy shows only a small performance advantage over pure MPI in some cases. Finally, improved implementations of two MPI collective operations on PC-SMP clusters are presented.

Shan, Hongzhang↗

The OpenMP Implementation of NAS Parallel Benchmarks and its Performance

As the new ccNUMA architecture became popular in recent years, parallel programming with compiler directives on these machines has evolved to accommodate new needs. In this study, we examine the effectiveness of OpenMP directives for parallelizing the NAS Parallel Benchmarks. Implementation details will be discussed and performance will be compared with the MPI implementation. We have demonstrated that OpenMP can achieve very good results for parallelization on a shared memory system, but effective use of memory and cache is very important.

Jin, Hao-Qiang↗

A software bus for thread objects

The authors have implemented a software bus for lightweight threads in an object-oriented programming environment that allows for rapid reconfiguration and reuse of thread objects in discrete-event simulation experiments. While previous research in object-oriented, parallel programming environments has focused on direct communication between threads, our lightweight software bus, called the MiniBus, provides a means to isolate threads from their contexts of execution by restricting communications between threads to message-passing via their local ports only. The software bus maintains a topology of connections between these ports. It routes, queues, and delivers messages according to this topology. This approach allows for rapid reconfiguration and reuse of thread objects in other systems without making changes to the specifications or source code. A layered approach that provides the needed transparency to developers is presented. Examples of using the MiniBus are given, and the value of bus architectures in building and conducting simulations of discrete-event systems is discussed.

Callahan, John R.↗

Parallel computing for probabilistic fatigue analysis

This paper presents the results of Phase I research to investigate the most effective parallel processing software strategies and hardware configurations for probabilistic structural analysis. We investigate the efficiency of both shared and distributed-memory architectures via a probabilistic fatigue life analysis problem. We also present a parallel programming approach, the virtual shared-memory paradigm, that is applicable across both types of hardware. Using this approach, problems can be solved on a variety of parallel configurations, including networks of single or multiprocessor workstations. We conclude that it is possible to effectively parallelize probabilistic fatigue analysis codes; however, special strategies will be needed to achieve large-scale parallelism to keep large number of processors busy and to treat problems with the large memory requirements encountered in practice. We also conclude that distributed-memory architecture is preferable to shared-memory for achieving large scale parallelism; however, in the future, the currently emerging hybrid-memory architectures will likely be optimal.

Sues, Robert H.↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗

ProperCAD: A portable object-oriented parallel environment for VLSI CAD

Most parallel algorithms for VLSI CAD proposed to date have one important drawback: they work efficiently only on machines that they were designed for. As a result, algorithms designed to date are dependent on the architecture for which they are developed and do not port easily to other parallel architectures. A new project under way to address this problem is described. A Portable object-oriented parallel environment for CAD algorithms (ProperCAD) is being developed. The objectives of this research are (1) to develop new parallel algorithms that run in a portable object-oriented environment (CAD algorithms using a general purpose platform for portable parallel programming called CARM is being developed and a C++ environment that is truly object-oriented and specialized for CAD applications is also being developed); and (2) to design the parallel algorithms around a good sequential algorithm with a well-defined parallel-sequential interface (permitting the parallel algorithm to benefit from future developments in sequential algorithms). One CAD application that has been implemented as part of the ProperCAD project, flat VLSI circuit extraction, is described. The algorithm, its implementation, and its performance on a range of parallel machines are discussed in detail. It currently runs on an Encore Multimax, a Sequent Symmetry, Intel iPSC/2 and i860 hypercubes, a NCUBE 2 hypercube, and a network of Sun Sparc workstations. Performance data for other applications that were developed are provided: namely test pattern generation for sequential circuits, parallel logic synthesis, and standard cell placement.

Ramkumar, Balkrishna↗

Parallel software support for computational structural mechanics

The application of the parallel programming methodology known as the Force was conducted. Two application issues were addressed. The first involves the efficiency of the implementation and its completeness in terms of satisfying the needs of other researchers implementing parallel algorithms. Support for, and interaction with, other Computational Structural Mechanics (CSM) researchers using the Force was the main issue, but some independent investigation of the Barrier construct, which is extremely important to overall performance, was also undertaken. Another efficiency issue which was addressed was that of relaxing the strong synchronization condition imposed on the self-scheduled parallel DO loop. The Force was extended by the addition of logical conditions to the cases of a parallel case construct and by the inclusion of a self-scheduled version of this construct. The second issue involved applying the Force to the parallelization of finite element codes such as those found in the NICE/SPAR testbed system. One of the more difficult problems encountered is the determination of what information in COMMON blocks is actually used outside of a subroutine and when a subroutine uses a COMMON block merely as scratch storage for internal temporary results.

Jordan, Harry F.↗

Space languages: Solving the classic scheduling problem in Ada and Lisp

The comparison of programming languages is best seen while evaluating similar systems. The strengths and weaknesses of both languages were investigated as the scheduler was being implemented. Some features used in both languages shall be object-oriented paradigms, parallel programming, search and production heuristics, and other classical artificial intelligence implementations.

Davis, Stephen↗

A portable MPI-based parallel vector template library

This paper discusses the design and implementation of a polymorphic collection library for distributed address-space parallel computers. The library provides a data-parallel programming model for C++ by providing three main components: a single generic collection class, generic algorithms over collections, and generic algebraic combining functions. Collection elements are the fourth component of a program written using the library and may be either of the built-in types of C or of user-defined types. Many ideas are borrowed from the Standard Template Library (STL) of C++, although a restricted programming model is proposed because of the distributed address-space memory model assumed. Whereas the STL provides standard collections and implementations of algorithms for uniprocessors, this paper advocates standardizing interfaces that may be customized for different parallel computers. Just as the STL attempts to increase programmer productivity through code reuse, a similar standard for parallel computers could provide programmers with a standard set of algorithms portable across many different architectures. The efficacy of this approach is verified by examining performance data collected from an initial implementation of the library running on an IBM SP-2 and an Intel Paragon.

Sheffler, Thomas J.↗

A Portable MPI-Based Parallel Vector Template Library

This paper discusses the design and implementation of a polymorphic collection library for distributed address-space parallel computers. The library provides a data-parallel programming model for C + + by providing three main components: a single generic collection class, generic algorithms over collections, and generic algebraic combining functions. Collection elements are the fourth component of a program written using the library and may be either of the built-in types of c or of user-defined types. Many ideas are borrowed from the Standard Template Library (STL) of C++, although a restricted programming model is proposed because of the distributed address-space memory model assumed. Whereas the STL provides standard collections and implementations of algorithms for uniprocessors, this paper advocates standardizing interfaces that may be customized for different parallel computers. Just as the STL attempts to increase programmer productivity through code reuse, a similar standard for parallel computers could provide programmers with a standard set of algorithms portable across many different architectures. The efficacy of this approach is verified by examining performance data collected from an initial implementation of the library running on an IBM SP-2 and an Intel Paragon.

Sheffler, Thomas J.↗

Astrometry of southern radio sources

An overview is presented of a number of astrometry and astrophysics programs based on radio sources from the Parkes 2.7 GHz catalogs. The programs cover the optical identification and spectroscopy of flat-spectrum Parkes sources and the determination of their milliarcsecond radio structures and positions. Work is also in progress to tie together the radio and Hipparcos positional reference frames. A parallel program of radio and optical astrometry of southern radio stars is also under way.

White, Graeme L.↗

An efficient data dependence analysis for parallelizing compilers

A novel algorithm, called the lambda test, is presented for an efficient and accurate data dependence analysis of multidimensional array references. It extends the numerical methods to allow all dimensions of array references to be tested simultaneously. Hence, it combines the efficiency and the accuracy of the both approaches. This algorithm has been implemented in PARAFRASE, a FORTRAN program parallelization restructurer developed at the University of Illinois at Urbana-Champaign. Some experimental results are presented to show its effectiveness.

Li, Zhiyuan↗

Fiats: Functional inference and training for surrogates

Fiats provides a platform for research on the training and deployment of neural-network surrogate models for computational science. Fiats also supports exploring, advancing, and combining functional, object-oriented, and parallel programming patterns in Fortran 2023. As such, the Fiats name has dual expansions: “Functional Inference And Training for Surrogates” or “Fortran Inference And Training for Science.” Fiats inference and training procedures are pure and therefore satisfy a language constraint imposed on procedure invocations inside Fortran’s parallel loop construct: do concurrent. Furthermore, the Fiats training procedures are built around a do concurrent parallel reduction. Several compilers can automatically parallelize do concurrent on Central Processing Units (CPUs) or Graphics Processing Units (GPUs). Fiats thus aims to achieve performance portability through standard language mechanisms.

Rouson, Damian [Lawrence Berkeley National Laborat↗

Parallelized reliability estimation of reconfigurable computer networks

A parallelized system, ASSURE, for computing the reliability of embedded avionics flight control systems which are able to reconfigure themselves in the event of failure is described. ASSURE accepts a grammar that describes a reliability semi-Markov state-space. From this it creates a parallel program that simultaneously generates and analyzes the state-space, placing upper and lower bounds on the probability of system failure. ASSURE is implemented on a 32-node Intel iPSC/860, and has achieved high processor efficiencies on real problems. Through a combination of improved algorithms, exploitation of parallelism, and use of an advanced microprocessor architecture, ASSURE has reduced the execution time on substantial problems by a factor of one thousand over previous workstation implementations. Furthermore, ASSURE's parallel execution rate on the iPSC/860 is an order of magnitude faster than its serial execution rate on a Cray-2 supercomputer. While dynamic load balancing is necessary for ASSURE's good performance, it is needed only infrequently; the particular method of load balancing used does not substantially affect performance.

Nicol, David M.↗

An integrated runtime and compile-time approach for parallelizing structured and block structured applications

Scientific and engineering applications often involve structured meshes. These meshes may be nested (for multigrid codes) and/or irregularly coupled (called multiblock or irregularly coupled regular mesh problems). A combined runtime and compile-time approach for parallelizing these applications on distributed memory parallel machines in an efficient and machine-independent fashion was described. A runtime library which can be used to port these applications on distributed memory machines was designed and implemented. The library is currently implemented on several different systems. To further ease the task of application programmers, methods were developed for integrating this runtime library with compilers for HPK-like parallel programming languages. How this runtime library was integrated with the Fortran 90D compiler being developed at Syracuse University is discussed. Experimental results to demonstrate the efficacy of our approach are presented. A multiblock Navier-Stokes solver template and a multigrid code were experimented with. Our experimental results show that our primitives have low runtime communication overheads. Further, the compiler parallelized codes perform within 20 percent of the code parallelized by manually inserting calls to the runtime library.

Agrawal, Gagan↗

Research on fission induced plasmas and nuclear pumped lasers at the Los Alamos Scientific Laboratory

A program of research on gaseous uranium and uranium plasmas is being conducted at The Los Alamos Scientific Laboratory under sponsorship of the National Aeronautics and Space Administration. The objective of this work is twofold: (1) to demonstrate the proof of principle of a gaseous uranium fueled reactor, and (2) pursue fundamental research on nuclear pumped lasers. The relevancy of the two parallel programs is embodied in the possibility of a high-performance uranium plasma reactor being used as the power supply for a nuclear pumped laser system. The accomplishments in the two above fields are summarized

Helmick, H. H.↗