Search NASA⌕ Search

SEARCH · Search NASA

Results for “application porting experiences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

OpenMP application experiences: Porting to accelerated nodes

As recent enhancements to the OpenMP specification become available in its implementations, there is a need to share the results of experimentation in order to better understand the OpenMP implementation’s behavior in practice, to identify pitfalls, and to learn how the implementations can be effectively deployed in scientific codes. We report on experiences gained and practices adopted when using OpenMP to port a variety of ECP applications, mini-apps and libraries based on different computational motifs to accelerator-based leadership-class high-performance supercomputer systems at the United States Department of Energy. Additionally, we identify important challenges and open problems related to the deployment of OpenMP. Through our report of experiences, we find that OpenMP implementations are successful on current supercomputing platforms and that OpenMP is a promising programming model to use for applications to be run on emerging and future platforms with accelerated nodes.

97 MATHEMATICS AND COMPUTING↗

Map Applications to Target Exascale Architecture with Machine-Specific Performance Analysis, Including Challenges and Projections

This Exascale Computing Project (ECP) milestone report summarizes the status of all 30 ECP Applications Development (AD) subprojects at the end of FY20. In October and November of 2020, a comprehensive assessment of AD projects was conducted by the ECP leadership. Reviews occurred virtually between October 27, 2020 and November 12, 2020. The review committee—consisting of the AD lead, deputy, and L3—was tasked with evaluating each subproject’s progress in porting their codes to early exascale architectures considered precursors to the planned exascale machines. This includes characterizing which modules have been ported to multi-accelerator nodes, initial performance analyses, the status of software integration, and a current vision of successes, obstacles, and next steps. As such, this report contains not only an accurate snapshot of each subproject’s current status but also represents an unprecedentedly broad account of experiences in porting large scientific applications to next-generation high-performance computing architectures.

97 MATHEMATICS AND COMPUTING↗

Lessons learned from an Ada conversion project

Recognizing the importance of building software with the robustness to accommodate rapid advances in technology, developers have focused on methods which preserve both past and future investment in software and in providing an advanced software engineering environment that frees the engineer, to the greatest extent possible, from the more routine aspects of software design and development. The software engineer is permitted to concentrate on the creative aspects of problem resolution. Standard languages such as Ada maximize portability across hardware and operating systems. Standard interfaces which enhance portability and permit the incorporation of new technology as it becomes available have been developed. Software design and development techniques which maximize portability receive increasing emphasis. Experience gained in porting an Ada application between two widely varying environments is evaluated in light of current practices to maximize software portability.

Porter, Tim↗

Ready for the Frontier: Preparing Applications for the World’s First Exascale System

Frontier, a supercomputer at the Oak Ridge Leadership Computing Facility (OLCF), debuted atop the Top500 list of the world’s most powerful supercomputers in June 2022 as the very first computer to produce exascale performance. Making sure scientific applications are optimized on this architecture is the critical link necessary to translate the newly available computational power into scientific insight and solutions. To that goal, the OLCF developed the Center for Accelerated Application Readiness (CAAR) program to ensure that a suite of highly optimized applications is ready for scientific runs at the onset of production operations for Frontier. This paper describes our experience in porting and optimizing such suite of applications in the OLCF’s CAAR program.

Budiardja, Reuben↗

Optimization and Portability of a Fusion OpenACC-based FORTRAN HPC Code from NVIDIA to AMD GPUs

NVIDIA has been the main provider of GPU hardware in HPC systems for over a decade. Most applications that benefit from GPUs have thus been developed and optimized for the NVIDIA software stack. Recent exascale HPC systems are, however, introducing GPUs from other vendors, e.g. with the AMD GPU-based OLCF Frontier system just becoming available. AMD GPUs cannot be directly accessed using the NVIDIA software stack, and require a porting effort by the application developers. This paper provides an overview of our experience porting and optimizing the CGYRO code, a widely-used fusion simulation tool based on FORTRAN with OpenACC-based GPU acceleration. While the porting from the NVIDIA compilers was relatively straightforward using the CRAY compilers on the AMD systems, the performance optimization required more fine-tuning. In the optimization effort, we uncovered code sections that had performed well on NVIDIA GPUs, but were unexpectedly slow on AMD GPUs. After AMD-targeted code optimizations, performance on AMD GPUs has increased to meet our expectations. Modest speed improvements were also seen on NVIDIA GPUs, which was an unexpected benefit of this exercise.

Sfiligoi, Igor↗

Application Experiences on a GPU-Accelerated Arm-based HPC Testbed

This paper assesses and reports the experience of ten teams working to port, validate, and benchmark several High Performance Computing applications on a novel GPU-accelerated Arm testbed system. The testbed consists of eight NVIDIA Arm HPC Developer Kit systems, each one equipped with a server-class Arm CPU from Ampere Computing and two data center GPUs from NVIDIA Corp. The systems are connected together using InfiniBand interconnect. The selected applications and mini-apps are written using several programming languages and use multiple accelerator-based programming models for GPUs such as CUDA, OpenACC, and OpenMP offloading. Working on application porting requires a robust and easy-to-access programming environment, including a variety of compilers and optimized scientific libraries. The goal of this work is to evaluate platform readiness and assess the effort required from developers to deploy well-established scientific workloads on current and future generation Arm-based GPU-accelerated HPC systems. The reported case studies demonstrate that the current level of maturity and diversity of software and tools is already adequate for large-scale production deployments.

Elwasif, Wael↗

Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part I)

This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.

Chapman, Barbara↗

Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part II)

This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.

Chapman, Barbara↗

Porting hypre to heterogeneous computer architectures: Strategies and experiences

We report that linear systems are occurring in many applications, and solving them can take a large amount of the total simulation time. The high performance library hypre provides a variety of interfaces and linear solvers, including various multigrid methods, that have achieved good scalability on a variety of homogeneous parallel computer architectures. Heterogeneous architectures with nodes that have both CPUs and accelerators provide new challenges, since they require more fine-grained parallelism and reduced data movement between different memories on a single node as well as across nodes. We will discuss our experiences and strategies to port hypre to heterogeneous computers with accelerators, including the design of a new memory model, the use of abstractions, the BoxLoop macros in the structured and semi-structured interfaces, and the restructuring of algebraic multigrid (AMG) into modular components. We present numerical experiments comparing CPU and GPU performance for several test problems.

97 MATHEMATICS AND COMPUTING↗

An Experimental Determination of Losses in a 3-Port Wave Rotor

Wave rotors, used in a gas turbine topping cycle, offer a potential route to higher specific power and lower specific fuel consumption. In order to exploit this potential properly, it is necessary to have some realistic means of calculating wave rotor performance, taking losses into account, so that wave rotors can be designed for good performance. This in turn requires a knowledge of the loss mechanisms. The experiment reported here was designed as a statistical experiment to identify the losses due to finite passage opening time, friction, and leakage. For simplicity, the experiment used a 3-port, flow divider, wave cycle, but the results should be applicable to other cycles. A 12 inch diameter rotor was used, with two different lengths, 9 inches and 18 inches, and two different passage widths, 0.25 inch and 0.54 inch, in order to vary friction and opening time. To vary leakage, moveable end-walls were provided so that the rotor to end-wall gap could be adjusted. The experiment is described, and the results are presented, together with a parametric fit to the data. The fit shows that there will be an optimum passage width for a given wave rotor, since, as the passage width increases, friction losses decrease, but opening-time losses increase, and vice-versa. Leakage losses can be made small at reasonable gap sizes.

Wilson, Jack↗

Design of the NASA Lewis 4-Port Wave Rotor Experiment

Pressure exchange wave rotors, used in a topping stage, are currently being considered as a possible means of increasing the specific power, and reducing the specific fuel consumption of gas turbine engines. Despite this interest, there is very little information on the performance of a wave rotor operating on the cycle (i.e., set of waves) appropriate for use in a topping stage. One such cycle, which has the advantage of being relatively easy to incorporate into an engine, is the four-port cycle. Consequently, an experiment to measure the performance of a four-port wave rotor for temperature ratios relevant to application as a topping cycle for a gas turbine engine has been designed and built at NASA Lewis. The design of the wave rotor is described, together with the constraints on the experiment.

Wilson, Jack↗

Early experiences evaluating the HPE/Cray ecosystem for AMD GPUs

Summary The Oak Ridge Leadership Computing Facility (OLCF) has a long history of supporting and promoting GPU‐accelerated computing starting with the deployment of the Titan supercomputer in 2021 and continuing with the Summit supercomputer which has a theoretical peak performance of approximately 200 petaflops. Because the majority of Summit's computational power comes from its 27,972 GPUs, users must port their applications to one of the supported programming models in order to make efficient use of the system. To prepare the transition to Frontier, the OLCF's exascale supercomputer, users will need to adapt to an entirely new ecosystem which will include new hardware and software technologies. First, users will need to familiarize themselves with the AMD Radeon GPU architecture. Furthermore, users who have been previously relying on CUDA will need to transition to the Heterogeneous‐Computing Interface for Portability (HIP) or one of the other supported programming models (e.g., OpenMP, OpenACC). In this work, we describe our initial experiences and lessons learned in porting three applications or proxy apps currently running on Summit to the HPE/Cray ecosystem to leverage the compute power from AMD GPUs: minisweep, GenASiS, and Sparkler. Each one is representative of current production workloads utilized at the OLCF, different programming languages, and different programming models.

Melesse Vergara, Verónica G.↗

Initial results from the NASA-Lewis wave rotor experiment

Wave rotors may play a role as topping cycles for jet engines, since by their use, the combustion temperature can be raised without increasing the turbine inlet temperature. In order to design a wave rotor for this, or any other application, knowledge of the loss mechanisms is required, and also how the design parameters affect those losses. At NASA LeRC, a 3-port wave rotor experiment operating on the flow-divider cycle, has been started with the objective of determining the losses. The experimental scheme is a three factor Box-Behnken design, with passage opening time, friction factor, and leakage gap as the factors. Variation of these factors is provided by using two rotors, of different length, two different passage widths for each rotor, and adjustable leakage gap. In the experiment, pressure transducers are mounted on the rotor, and give pressure traces as a function of rotational angle at the entrance and exit of a rotor passage. In addition, pitot rakes monitor the stagnation pressures for each port, and orifice meters measure the mass flows. The results show that leakage losses are very significant in the present experiment, but can be reduced considerably by decreasing the rotor to wall clearance spacing.

Wilson, Jack↗

Initial results from the NASA Lewis wave rotor experiment

Wave rotors may play a role as topping cycles for jet engines, since by their use, the combustion temperature can be raised without increasing the turbine inlet temperature. In order to design a wave rotor for this, or any other application, knowledge of the loss mechanisms is required, and also how the design parameters affect those losses. At NASA LeRC, a 3-port wave rotor experiment operating on the flow-divider cycle, has been started with the objective of determining the losses. The experimental scheme is a three factor Box-Behnken design, with passage opening time, friction factor, and leakage gap as the factors. Variation of these factors is provided by using two rotors, of different length, two different passage widths for each rotor, and adjustable leakage gap. In the experiment, pressure transducers are mounted on the rotor, and give pressure traces as a function of rotational angle at the entrance and exit of a rotor passage. In addition, pitot rakes monitor the stagnation pressures for each port, and orifice meters measure the mass flows. The results show that leakage losses are very significant in the present experiment, but can be reduced considerably by decreasing the rotor to wall clearance spacing.

Wilson, Jack↗

Parallelization of NAS Benchmarks for Shared Memory Multiprocessors

This paper presents our experiences of parallelizing the sequential implementation of NAS benchmarks using compiler directives on SGI Origin2000 distributed shared memory (DSM) system. Porting existing applications to new high performance parallel and distributed computing platforms is a challenging task. Ideally, a user develops a sequential version of the application, leaving the task of porting to new generations of high performance computing systems to parallelization tools and compilers. Due to the simplicity of programming shared-memory multiprocessors, compiler developers have provided various facilities to allow the users to exploit parallelism. Native compilers on SGI Origin2000 support multiprocessing directives to allow users to exploit loop-level parallelism in their programs. Additionally, supporting tools can accomplish this process automatically and present the results of parallelization to the users. We experimented with these compiler directives and supporting tools by parallelizing sequential implementation of NAS benchmarks. Results reported in this paper indicate that with minimal effort, the performance gain is comparable with the hand-parallelized, carefully optimized, message-passing implementations of the same benchmarks.

Waheed, Abdul↗

NAS Experiences of Porting CM Fortran Codes to HPF on IBM SP2 and SGI Power Challenge

Current Connection Machine (CM) Fortran codes developed for the CM-2 and the CM-5 represent an important class of parallel applications. Several users have employed CM Fortran codes in production mode on the CM-2 and the CM-5 for the last five to six years, constituting a heavy investment in terms of cost and time. With Thinking Machines Corporation's decision to withdraw from the hardware business and with the decommissioning of many CM-2 and CM-5 machines, the best way to protect the substantial investment in CM Fortran codes is to port the codes to High Performance Fortran (HPF) on highly parallel systems. HPF is very similar to CM Fortran and thus represents a natural transition. Conversion issues involved in porting CM Fortran codes on the CM-5 to HPF are presented. In particular, the differences between data distribution directives and the CM Fortran Utility Routines Library, as well as the equivalent functionality in the HPF Library are discussed. Several CM Fortran codes (Cannon algorithm for matrix-matrix multiplication, Linear solver Ax=b, 1-D convolution for 2-D datasets, Laplace's Equation solver, and Direct Simulation Monte Carlo (DSMC) codes have been ported to Subset HPF on the IBM SP2 and the SGI Power Challenge. Speedup ratios versus number of processors for the Linear solver and DSMC code are presented.

Saini, Subhash↗