Search NASA⌕ Search

SEARCH · Search NASA

Results for “application porting experiences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

OpenMP application experiences: Porting to accelerated nodes

As recent enhancements to the OpenMP specification become available in its implementations, there is a need to share the results of experimentation in order to better understand the OpenMP implementation’s behavior in practice, to identify pitfalls, and to learn how the implementations can be effectively deployed in scientific codes. We report on experiences gained and practices adopted when using OpenMP to port a variety of ECP applications, mini-apps and libraries based on different computational motifs to accelerator-based leadership-class high-performance supercomputer systems at the United States Department of Energy. Additionally, we identify important challenges and open problems related to the deployment of OpenMP. Through our report of experiences, we find that OpenMP implementations are successful on current supercomputing platforms and that OpenMP is a promising programming model to use for applications to be run on emerging and future platforms with accelerated nodes.

97 MATHEMATICS AND COMPUTING↗

Map Applications to Target Exascale Architecture with Machine-Specific Performance Analysis, Including Challenges and Projections

This Exascale Computing Project (ECP) milestone report summarizes the status of all 30 ECP Applications Development (AD) subprojects at the end of FY20. In October and November of 2020, a comprehensive assessment of AD projects was conducted by the ECP leadership. Reviews occurred virtually between October 27, 2020 and November 12, 2020. The review committee—consisting of the AD lead, deputy, and L3—was tasked with evaluating each subproject’s progress in porting their codes to early exascale architectures considered precursors to the planned exascale machines. This includes characterizing which modules have been ported to multi-accelerator nodes, initial performance analyses, the status of software integration, and a current vision of successes, obstacles, and next steps. As such, this report contains not only an accurate snapshot of each subproject’s current status but also represents an unprecedentedly broad account of experiences in porting large scientific applications to next-generation high-performance computing architectures.

97 MATHEMATICS AND COMPUTING↗

Ready for the Frontier: Preparing Applications for the World’s First Exascale System

Frontier, a supercomputer at the Oak Ridge Leadership Computing Facility (OLCF), debuted atop the Top500 list of the world’s most powerful supercomputers in June 2022 as the very first computer to produce exascale performance. Making sure scientific applications are optimized on this architecture is the critical link necessary to translate the newly available computational power into scientific insight and solutions. To that goal, the OLCF developed the Center for Accelerated Application Readiness (CAAR) program to ensure that a suite of highly optimized applications is ready for scientific runs at the onset of production operations for Frontier. This paper describes our experience in porting and optimizing such suite of applications in the OLCF’s CAAR program.

Budiardja, Reuben↗

Optimization and Portability of a Fusion OpenACC-based FORTRAN HPC Code from NVIDIA to AMD GPUs

NVIDIA has been the main provider of GPU hardware in HPC systems for over a decade. Most applications that benefit from GPUs have thus been developed and optimized for the NVIDIA software stack. Recent exascale HPC systems are, however, introducing GPUs from other vendors, e.g. with the AMD GPU-based OLCF Frontier system just becoming available. AMD GPUs cannot be directly accessed using the NVIDIA software stack, and require a porting effort by the application developers. This paper provides an overview of our experience porting and optimizing the CGYRO code, a widely-used fusion simulation tool based on FORTRAN with OpenACC-based GPU acceleration. While the porting from the NVIDIA compilers was relatively straightforward using the CRAY compilers on the AMD systems, the performance optimization required more fine-tuning. In the optimization effort, we uncovered code sections that had performed well on NVIDIA GPUs, but were unexpectedly slow on AMD GPUs. After AMD-targeted code optimizations, performance on AMD GPUs has increased to meet our expectations. Modest speed improvements were also seen on NVIDIA GPUs, which was an unexpected benefit of this exercise.

Sfiligoi, Igor↗

Application Experiences on a GPU-Accelerated Arm-based HPC Testbed

This paper assesses and reports the experience of ten teams working to port, validate, and benchmark several High Performance Computing applications on a novel GPU-accelerated Arm testbed system. The testbed consists of eight NVIDIA Arm HPC Developer Kit systems, each one equipped with a server-class Arm CPU from Ampere Computing and two data center GPUs from NVIDIA Corp. The systems are connected together using InfiniBand interconnect. The selected applications and mini-apps are written using several programming languages and use multiple accelerator-based programming models for GPUs such as CUDA, OpenACC, and OpenMP offloading. Working on application porting requires a robust and easy-to-access programming environment, including a variety of compilers and optimized scientific libraries. The goal of this work is to evaluate platform readiness and assess the effort required from developers to deploy well-established scientific workloads on current and future generation Arm-based GPU-accelerated HPC systems. The reported case studies demonstrate that the current level of maturity and diversity of software and tools is already adequate for large-scale production deployments.

Elwasif, Wael↗

Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part I)

This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.

Chapman, Barbara↗

Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part II)

This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.

Chapman, Barbara↗

Porting hypre to heterogeneous computer architectures: Strategies and experiences

We report that linear systems are occurring in many applications, and solving them can take a large amount of the total simulation time. The high performance library hypre provides a variety of interfaces and linear solvers, including various multigrid methods, that have achieved good scalability on a variety of homogeneous parallel computer architectures. Heterogeneous architectures with nodes that have both CPUs and accelerators provide new challenges, since they require more fine-grained parallelism and reduced data movement between different memories on a single node as well as across nodes. We will discuss our experiences and strategies to port hypre to heterogeneous computers with accelerators, including the design of a new memory model, the use of abstractions, the BoxLoop macros in the structured and semi-structured interfaces, and the restructuring of algebraic multigrid (AMG) into modular components. We present numerical experiments comparing CPU and GPU performance for several test problems.

97 MATHEMATICS AND COMPUTING↗

Early experiences evaluating the HPE/Cray ecosystem for AMD GPUs

Summary The Oak Ridge Leadership Computing Facility (OLCF) has a long history of supporting and promoting GPU‐accelerated computing starting with the deployment of the Titan supercomputer in 2021 and continuing with the Summit supercomputer which has a theoretical peak performance of approximately 200 petaflops. Because the majority of Summit's computational power comes from its 27,972 GPUs, users must port their applications to one of the supported programming models in order to make efficient use of the system. To prepare the transition to Frontier, the OLCF's exascale supercomputer, users will need to adapt to an entirely new ecosystem which will include new hardware and software technologies. First, users will need to familiarize themselves with the AMD Radeon GPU architecture. Furthermore, users who have been previously relying on CUDA will need to transition to the Heterogeneous‐Computing Interface for Portability (HIP) or one of the other supported programming models (e.g., OpenMP, OpenACC). In this work, we describe our initial experiences and lessons learned in porting three applications or proxy apps currently running on Summit to the HPE/Cray ecosystem to leverage the compute power from AMD GPUs: minisweep, GenASiS, and Sparkler. Each one is representative of current production workloads utilized at the OLCF, different programming languages, and different programming models.

Melesse Vergara, Verónica G.↗

Imaging Bragg Edge Analysis TooLs for Engineering Structures (iBeatles)

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory (ORNL) provides pulsed neutrons with energies varying from epithermal to cold. In preparation for VENUS, the neutron imaging beamline to be located at beam port 10, we have performed a series of experiments focused on wavelength-dependent radiography and computed tomography for a broad range of applications, from materials science to biological tissues.One of the time-of-flight (TOF) techniques that is of interest to the scientific community is the 2-dimensional mapping of phases and average crystalline plane orientation in samples both ex-situ and during applied stresses such as tensile loading and heating. This technique is known as Bragg edgeimaging and relies on the identification of changes of transmission values, fitting of the edge to measure its displacement, and thus identify the shift in lattice parameter due to stresses. One of the challenges of TOF imaging measurements is the amount of data and the inability to observe Bragg edge shifts in real time during an experiment. Thus, we have been focusing on creating a Python-based interface that allows fast data processing and instantaneous mapping and fitting of the Bragg edges, and their evolution through time. Python libraries and Jupyter notebooks have been implemented to facilitate decision making during an experiment. The advantage of the notebooks is the possibility to guide an experiment as they can quickly process and display Bragg edge data. These notebooks can be used independently, or can be combined in a Python Graphical User Interface (GUI) tool called iBeatles. This interface permits visualization and fitting of the Bragg edges, and ultimately back-projects the fitting results onto the radiographs to display a strain map. Assuming data collection has sufficient statistics, the strain mapping analysis can be performed on a pixel-by-pixel basis. This development is a step forward toward a better user experience at the future VENUS beamline in terms of live feedback and productivity. Analysis that used to take days of switching between different applications can now be done in minutes within the

Bilheux, JeanChristophe [Oak Ridge National Labora↗

Programming approaches for scalability, performance, and portability of combustion physics codes

Here, this paper presents the process, strategy, and results associated with porting a typical combustion physics flow solver to current state-of-the-art and future massively-parallel computer architectures. Major focus is placed on the distinct algorithmic structure of these types of codes and how it can be integrated with modern programming paradigms for heterogeneous platforms (i.e., distributed many-core systems with accelerators). An end-to-end case study is presented that exemplifies the process in a generic manner, which then serves as a clear guide with respect to the strategy and best practices leading to a robust and adaptable framework that performs well, is durable over time, is portable, and requires minimal human-effort. This end is accomplished beginning with the use of a mature, validated, structured, multiblock code framework optimized for application of both Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS). This code has been ported to a variety of platforms over the past decade, including most recently the Oak Ridge Leadership Computing Facility’s “Summit” Platform. The experience gained on these multiple platforms provides general insights and thus the results presented are not specific to any one code or platform other than the overarching trend toward distributed many-core systems with accelerators in order to move toward exascale performance. The resultant performance and scalability of the ported code is demonstrated on a real-world application; a state-of-the-art rotating detonation rocket engine simulation that matches the complex geometry and boundary conditions imposed as part of a companion experimental campaign.

97 MATHEMATICS AND COMPUTING↗

Experiences with Porting the FLASH Code to Ookami, an HPE Apollo 80 A64FX Platform

We present initial experiences with running the community simulation code FLASH, developed at the University of Chicago for multi-scale multi-physics applications, on Ookami, a technology testbed featuring the A64FX processor developed by Fujitsu. Our effort focused largely on running FLASH “right out of the box” to see which combinations of compilers and software implementations (e.g. MPI) allowed the code to run with minimal modification. FLASH was one application in a larger effort to deploy Ookami; it served as a test for different versions of newly installed software, and as a cornerstone for the FAQ page of the Ookami website. Here, we report on our results with different compilers and other software, along with our initial scaling results and attempts to utilize the A64FX’s SVE instructions and NUMA architecture. We found that FLASH readily ran with different compilers and MPI implementations, and showed the expected good scaling with no turning. However, more work must be done to fully take advantage of the A64FX’s architectural features and produce a significant speedup for FLASH on Ookami.

79 ASTRONOMY AND ASTROPHYSICS↗

7.2 kV Three-Port SiC Single-Stage Current-Source Solid-State Transformer With 90 kV Lightning Protection

This article proposes a multiport modular single-stage current-source solid-state transformer (SST) for applications like photovoltaic, energy storage integration, electric vehicle fast charging, data center, etc. The 7.2 kV 50 kVA current-source SST consists of five input-series output-parallel modules, each based on 3.3 kV SiC reverse-blocking MOSFET-plus-diode modules. The proposed SST has some unique features. First, compared to the voltage-source or matrix converter-based SSTs, the current-source SST has a unique advantage of single-stage isolated AC/DC or AC/AC conversion with an inductive DC link, but no medium-voltage (MV) AC experiments have been reported. This article for the first time demonstrates MV AC current-source SST up to 7.5 kV peak. Second, the multiport SST has a buffer port for active power decoupling (APD) or energy storage integration. The double-line-frequency power ripple from single-phase AC grid normally results in a large capacitor size in MV SSTs. The APD scheme is proposed in MV applications for the first time to enable a reduced DC link and the electrolytic capacitor-less SST with high reliability. Third, as a direct grid-connected converter without line-frequency transformer, insulation and protection are critical. A medium-frequency transformer design passes 55 kV basic-insulation level (BIL) and 60 kV high potential dielectrics withstand test with only 0.09% leakage inductance. Importantly, a lightning protection scheme is presented to protect the SST itself from 90 kV BIL impulse. Fourth, the proposed current-source SST topology is a modular soft-switching solid-state transformer (M-S4T) with full-range zero-voltage switching and controlled dv/dt for low electromagnetic interference. Furthermore, these concepts are verified in a three-port M-S4T prototype with forced oil cooling under single-module, stacked-module, steady-state, and dynamic operations.

14 SOLAR ENERGY↗

Development and Demonstration of a Prototype Molten Salt Sampling System

Molten salt reactors (MSRs) offer potential operability and safety advantages when compared to commercial light water reactors (LWRs). However, operating experience with MSRs is sparse in comparison to what exists for LWRs. Further, the chemical and isotopic composition of the fuel and/or coolant salt is dynamic and difficult to characterize continuously, posing potential safety, operability, and safeguards unknowns that need to be addressed. A molten salt sampling system (MSSS) is regarded as a necessary subsystem within first generation MSRs used to obtain samples of salt for chemical and isotopic analysis in support of the need to monitor and control salt composition during operation. The MSSS is being developed using the Safety-in-Design (SiD) methodology, which incorporates incremental integration of safety analysis into the design process. The MSSS conceptual design emerging from the application of the early stages of the SiD methodology consists of a sample collection system and its housing, a freeze port, and inert gas control and delivery systems. This article describes the prototypes developed to test the functions of these MSSS subsystems, presents the results of testing in both dry and molten salt environments (including reliability data collection performed in accordance with the principles of SiD and the development of a semiquantitative fault tree model), and summarizes the opportunities for future design and testing enhancements based on the results of prototype testing.

molten salt reactor↗

Standardized Protocol for Real-Time APIs as Required by Title 23 CFR 680.116(c)

Improving the ability of drivers to easily locate working and available chargers is key to improving the public charging experience. Electric vehicle charging providers who are recipients of federal funds through the National Electric Vehicle Infrastructure (NEVI) Formula Program, Charging and Fueling Infrastructure (CFI) Discretionary Grant Program, and other funding programs as identified under Title 23 of the U.S. Code must deploy and maintain an application programming interface (API) to access information about charging stations they operate.1 This includes information about individual charging ports, pricing, and availability in accordance with the Federal Highway Administration’s National Electric Vehicle Infrastructure Standards and Requirements, 23 CFR 680.116(c), herein referred to as the minimum standards (Federal Highway Administration 2023). Specifically outlined in the minimum standards, states and other designated recipients are required to ensure that charging station information including location, connector type, power level, real-time status, and real-time price to charge are available free of charge to third-party software developers through an API. These requirements are intended to enable effective communication with consumers about available charging stations and help consumers make informed decisions about trip planning, including when and where to charge. This document provides a standardized protocol for how to structure data, data update frequency, and practices for making the data required to be shared via API usable for improving public transparency and the customer experience. These are recommendations only and do not modify the Federal Highway Administration’s minimum standards in any way.

33 ADVANCED PROPULSION SYSTEMS↗

Optode Images, and Dissolved Oxygen Data from Sand and Sediment Flow-through Column Experiments associated with: "2-D Imaging of Dissolved Oxygen Concentration in Flow-Through Sediment Columns"

This data package is associated with the publication "2-D Imaging of Dissolved Oxygen Concentration in Flow-Through Sediment Columns" submitted to Frontiers in Water (Garayburu-Caruso et al., 2022). This work presents a novel application of planar optode technology to sediment columns that allows the study of well-constrained flow fields without the need to assume completely homogeneous flow. We introduce a flow-through column with the interior coated in an oxygen sensing film and a series of sampling ports as a unique tool capable of capturing changes in dissolved oxygen (DO) at high frequency (10 s) and spatial (< 1 mm) resolution across the length of a column. In addition, we provide custom scripts that allows image processing steps to be automated and consistently reproduced over a variety of experimental conditions. This dataset is comprised of four folders (1) 01_Riverbed_Sediment_Experiment, (2) 02_Sand_Experiment, (3) 01_Riverbed_Sediment_Raw_Images, and (4) 02_Sand_Experiment_Raw_Images. 01_Riverbed_Sediment_Experiment contains (1) a subfolder with csv files that contain statistical parameters from the linear regressions applied to each optode image, (2) a subfolder with csv files from DO data measured by in-line sensors, (3) a subfolder with csv files containing column and injection specific input parameters for the image processing script, and (4) a subfolder with scripts used to process the images and create associated figures. 02_Sand_Experiment contains (1) a subfolder with csv files containing column and injection-specific input parameters for the image processing script, and (2) a subfolder with scripts used to process the images and create associated figures. 01_Riverbed_Sediment_Raw_Images (spited in 3 parts due to size) contains a subfolder with raw red, and green images from the optode system during the MilliQ Injection, Vanillin Injection or Vanillin + Sampling Injection respectively. 02_Sand_Experiment_Raw_Images contains a subfolder with raw red, and green images from the optode system. Outside of the main folders there is a csv containing file-level metadata and a csv data dictionary defining column headers for all csv files contained in the data package.

54 ENVIRONMENTAL SCIENCES↗