Search NASA⌕ Search

SEARCH · Search NASA

Results for “Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

JACC.shared: Leveraging HPC Metaprogramming and Performance Portability for Computations That Use Shared Memory GPUs

In this work, we present JACC.shared, a new feature of Julia for ACCelerators (JACC), which is the performanceportable and metaprogramming model of the just-in-time and LLVM-based Julia language. This new feature allows JACC applications to leverage the high-performance computing (HPC) capabilities of high-bandwidth, on-chip GPU memory. Historically, exploiting high-bandwidth, shared-memory GPUs has not been a priority for high-level programming solutions. JACC.shared covers that gap for the first time, thereby providing a highlevel, portable, and easy-to-use solution for programmers to exploit this memory and supporting all current major accelerator architectures. Well-known HPC and AI workloads, such as multi/hyperspectral imaging and AI convolutions, have been used to evaluate JACC.shared on two exascale GPU architectures hosted by some of the most powerful US Department of Energy supercomputers: Perlmutter (NVIDIA A100) and Frontier (AMD MI250X). The performance evaluation reports speedup of up to 3.5× by adding only one line of code to the base codes, thus providing important accelerators in a simple, portable, and transparent way and elevating the programming productivity and performance-portability capabilities for Julia/JACC HPC, AI, and scientific applications.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

JACC: Leveraging HPC Meta-Programming and Performance Portability with the Just-in-Time and LLVM-based Julia Language

We present JACC (Julia for Accelerators), the first high-level, and performance-portable model for the just-in-time and LLVM-based Julia language. JACC provides a unified and lightweight front end across different back ends available in Julia, enabling the same Julia code to run efficiently on many HPC CPU and GPU targets. We evaluated the performance of JACC for common HPC kernels as well as for the most computationally demanding kernels used in applications, HPCCG, a supercomputing benchmark test for sparse domains, and HARVEY, a blood flow simulator to assist in the diagnosis and treatment of patients suffering from vascular diseases. We carried out the performance analysis on the most advanced US DOE supercomputers: Aurora, Frontier, and Perlmutter. Overall, we show that JACC has a negligible overhead versus vendor-specific solutions, reporting GPU speedups with no extra cost to programmability.

Valero-Lara, Pedro↗

Portable EDITOR (PEDITOR): A portable image processing system

The PEDITOR image processing system was created to be readily transferable from one type of computer system to another. While nearly identical in function and operation to its predecessor, EDITOR, PEDITOR employs additional techniques which greatly enhance its portability. These cover system structure and processing. In order to confirm the portability of the software system, two different types of computer systems running greatly differing operating systems were used as target machines. A DEC-20 computer running the TOPS-20 operating system and using a Pascal Compiler was utilized for initial code development. The remaining programmers used a Motorola Corporation 68000-based Forward Technology FT-3000 supermicrocomputer running the UNIX-based XENIX operating system and using the Silicon Valley Software Pascal compiler and the XENIX C compiler for their initial code development.

Angelici, G.↗

Portable programming on parallel/networked computers using the Application Portable Parallel Library (APPL)

The Application Portable Parallel Library (APPL) is a subroutine-based library of communication primitives that is callable from applications written in FORTRAN or C. APPL provides a consistent programmer interface to a variety of distributed and shared-memory multiprocessor MIMD machines. The objective of APPL is to minimize the effort required to move parallel applications from one machine to another, or to a network of homogeneous machines. APPL encompasses many of the message-passing primitives that are currently available on commercial multiprocessor systems. This paper describes APPL (version 2.3.1) and its usage, reports the status of the APPL project, and indicates possible directions for the future. Several applications using APPL are discussed, as well as performance and overhead results.

Quealy, Angela↗

Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ Workflow

In-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms.

Zhang, Bo [University of Utah]↗

Lessons Learned From the Construction of a Portable Cleanroom for NASA OSIRIS-REx Mission Deintegration

NASA Johnson Space Center (JSC) Infrastructure and Astromaterials Acquisition & Curation Office completed construction and commissioning of the OSIRIS-REx (OREx) Deintegration portable cleanroom at the Utah Test and Training Range (UTTR). The new portable cleanroom was designed to receive the OREx sample return capsule from the landing point on the range to an ISO7 environment. Scientists used the portable clean-room to deintegrate the sample canister from the sample return capsule. Once separated, the sample canister was put in a container under nitrogen purge for transportation to B31 at the Johnson Space Center for astromaterial sample extraction, preliminary analysis, and long-term curation. The portable cleanroom was built by a subcontractor at their facility and then deconstructed to be transported to the remote location at UTTR. Since construction was completed in a remote location all tools and materials had to be transported from contractor site in Dallas, TX. The cleanroom was constructed within an existing facility, which provided conditioned air, electric power, and protection from the elements. Careful coordination was required between the host facility, cleanroom contractor, mission scientists, and JSC facilities and curation personnel. An existing anteroom at JSC was transported to UTTR and added to the portable cleanroom after there was concern about contamination without one for personnel entry/exit. The scientific study of organics is critical for the mission, so a stringent contamination control plan was implemented for low organics. Given these mission requirements the cleanroom construction materials were carefully selected to not hinder the scientific search for amino acids and the study of organics in the samples. The same cleanroom contractor that built the long-term astromaterial curation cleanroom back at JSC Houston, TX was selected to build the portable cleanroom and instructed to use the same materials. The cleanroom had double doors to open and allow the sample return capsule to fit into the cleanroom on its stand and be transferred to a clean stand already in the cleanroom. The portable cleanroom successfully completed its mission and the sample canister was safely deintegrated and transported to JSC under nitrogen purge.

astromaterials curation↗

Portable Cleanroom for NASA OSIRIS-REx Mission Deintegration

NASA Johnson Space Center (JSC) Infrastructure and Astromaterials Acquisition & Curation Office completed construction and commissioning of the OSIRIS-REx (OREx) Deintegration portable cleanroom at the Utah Test and Training Range (UTTR). The new portable cleanroom was designed to receive the OREx sample return capsule from the landing point on the range to an ISO7 environment. Scientists used the portable clean-room to deintegrate the sample canister from the sample return capsule. Once separated, the sample canister was put in a container under nitrogen purge for transportation to B31 at the Johnson Space Center for astromaterial sample extraction, preliminary analysis, and long-term curation. The portable cleanroom was built by a subcontractor at their facility and then deconstructed to be transported to the remote location at UTTR. Since construction was completed in a remote location all tools and materials had to be transported from contractor site in Dallas, TX. The cleanroom was constructed within an existing facility, which provided conditioned air, electric power, and protection from the elements. Careful coordination was required between the host facility, cleanroom contractor, mission scientists, and JSC facilities and curation personnel. An existing anteroom at JSC was transported to UTTR and added to the portable cleanroom after there was concern about contamination without one for personnel entry/exit. The scientific study of organics is critical for the mission, so a stringent contamination control plan was implemented for low organics. Given these mission requirements the cleanroom construction materials were carefully selected to not hinder the scientific search for amino acids and the study of organics in the samples. The same cleanroom contractor that built the long-term astromaterial curation cleanroom back at JSC Houston, TX was selected to build the portable cleanroom and instructed to use the same materials. The cleanroom had double doors to open and allow the sample return capsule to fit into the cleanroom on its stand and be transferred to a clean stand already in the cleanroom. The portable cleanroom successfully completed its mission and the sample canister was safely deintegrated and transported to JSC under nitrogen purge.

astromaterials curation↗

CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous Systems

Performance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems.

Fujita, Norihisa↗

Assessing the Added Value of Miniature X-Ray in the Setting of Portable Ultrasound in Spaceflight

INTRODUCTION: Point-of-care ultrasound (POCUS) has become the standard of care for imaging diagnosis and management in low-Earth Orbit (LEO) spaceflight and it has long been hypothesized that POCUS will also be the standard of care for exploration spaceflight. However, like the trajectory for which ultrasound became more portable and user-friendly, the mass, volume, and power requirements of radiography devices for both diagnostic and therapeutic applications have also been dramatically reduced. This study seeks to determine the clinical utility and added value of miniature x-ray (XR) for the diagnosis and management of each of the 119 conditions within NASA Exploration Medical Capability’s IMPACT Condition List (ICL) given that the medical system is presumed to already be carrying a handheld portable POCUS device. METHODS: For each condition, a team of reviewers performed a rapid systematic literature review seeking sensitivity and specificity data for both handheld portable ultrasound and miniature XR. When there was a paucity of data, subject matter expertise and clinical experience was added to semi-quantitatively score the added value of miniature XR, given an US was already available for both diagnosis and management. Diagnostic utility of a modality for a condition was evaluated in the setting of both the best- and worst-case scenario definitions included and defined by the ICL. Evidence tracing and quality of evidence scores were also recorded. RESULTS/DISCUSSION: Conditions for which it was determined that miniature XR added diagnostic or therapeutic value are provided in this presentation. Previously presented work by our team has demonstrated that XR provides diagnostic and management capabilities that are hypothesized to complement or surpass ultrasound for over one-third of medical conditions that may arise during exploration spaceflight (i.e., diagnosis of injuries to the axial skeleton, teeth, and lungs as well as management of orthopedic reductions, endotracheal tube placement, and drain placement confirmation). In the setting of known inclusion of a handheld portable POCUS device, there remains significant added value of portable miniature XR. Whether or not this added clinical benefit is worth the mass, volume, and power requirements of the radiography system remains yet unknown and is the focus of future work. LEARNING OBJECTIVES: 1) Understand the value of ultrasound and radiography in the diagnosis and management of medical comorbidities that may arise in exploration spaceflight; 2) Understand the medical conditions of highest concern on exploration class missions for which miniature x-ray may provide added value to portable ultrasound.

J G Steller↗

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)↗

Enabling Scientific Applications with Performance-Portability and High-Productivity for Multi-GPU Programming with JACC.Multi

This work bridges the gap between multi-GPU computing and high-productivity, performance-portable programming solutions. Our goal is to enhance scientific applications with a productive and portable solution—program once, deploy everywhere—for multi-GPU programming with no cost to programmability. To accomplish this, we implemented JACC.Multi, which is part of the Julia for ACCelerators (JACC) performance-portable framework. JACC. Multi is the only high-level, portable metaprogramming solution that targets multi-GPU environments and is integrated in a readily accessible programming language (e.g., Julia language). With transparent GPU-to-GPU communication, JACC. Multi is optimized for scientific application workloads and is portable for NVIDIA and AMD accelerators. For the evaluation, we use two modern multi-GPU systems: Hudson, which features two NVIDIA H100 Hopper GPUs per node, and Frontier, which features four AMD MI250X GPUs per node, each with two Graphics Compute Dies (GCDs) for a total of eight GCDs per node. Additionally, as part of the evaluation, we use JACC (one GPU), MPI+JACC, and JACC. Multi codes that implement well-known and widely used scientific algorithms/kernels such as the conjugate gradient algorithm and an explicit forward Euler solver that requires GPU-to-GPU communication. Overall, JACC. Multi codes achieve better performance than MPI+JACC codes and significant speedups over JACC (one GPU), with up to 1.9× on Hudson and 6× on Frontier.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Precise time dissemination via portable atomic clocks

The most precise operational method of time dissemination over long distances presently available to the Precise Time and Time Interval (PTTI) community of users is by means of portable atomic clocks. The Global Positioning System (GPS), the latest system showing promise of replacing portable clocks for global PTTI dissemination, was evaluated. Although GPS has the technical capability of providing superior world-wide dissemination, the question of present cost and future accessibility may require a continued reliance on portable clocks for a number of years. For these reasons a study of portable clock operations as they are carried out today was made. The portable clock system that was utilized by the U.S. Naval Observatory (NAVOBSY) in the global synchronization of clocks over the past 17 years is described and the concepts on which it is based are explained. Some of its capabilities and limitations are also discussed.

Putkovich, K.↗

Satellite sound broadcasting system, portable reception

Studies are underway at JPL in the emerging area of Satellite Sound Broadcast Service (SSBS) for direct reception by low cost portable, semi portable, mobile and fixed radio receivers. This paper addresses the portable reception of digital broadcasting of monophonic audio with source material band limited to 5 KHz (source audio comparable to commercial AM broadcasting). The proposed system provides transmission robustness, uniformity of performance over the coverage area and excellent frequency reuse. Propagation problems associated with indoor portable reception are considered in detail and innovative antenna concepts are suggested to mitigate these problems. It is shown that, with the marriage of proper technologies a single medium power satellite can provide substantial direct satellite audio broadcast capability to CONUS in UHF or L Bands, for high quality portable indoor reception by low cost radio receivers.

Golshan, Nasser↗

Evaluating Application Characteristics for GPU Portability Layer Selection

GPUs have become the dominant source of computing power for high performance computing and are increasingly being used across the High Energy Physics computing landscape for a wide variety of tasks. Though NVIDIA is currently the main provider of GPUs, AMD and Intel are rapidly increasing their market share. As a result, programming using a vendor-specific language such as CUDA can significantly reduce deployment choices. There are a number of portability layers such as Kokkos, Alpaka, SYCL, OpenMP and std::par that permit execution on a broad range of GPU and CPU architectures, significantly increasing the flexibility of application programmers. However, each of these portability layers has its own characteristics, performing better at some tasks and worse at others, or placing limitations on aspects of the application. In this presentation, we report on a study of application and kernel characteristics that can influence the choice of a portability layer and show how each layer handles these characteristics. We have analyzed representative heterogeneous applications from CMS (patatrack and p2r), DUNE (Wire-Cell Toolkit), and ATLAS (FastCaloSim) to identify key application characteristics that have different behaviors for the various portability technologies. Using these results, developers can make more informed decisions on which GPU portability technology is best suited to their application.

Atif, Mohammad [Brookhaven]↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

On a Simplified Approach to Achieve Parallel Performance and Portability Across CPU and GPU Architectures

This paper presents software advances to easily exploit computer architectures consisting of a multi-core CPU and CPU+GPU to accelerate diverse types of high-performance computing (HPC) applications using a single code implementation. The paper describes and demonstrates the performance of the open-source C++ matrix and array (MATAR) library that uniquely offers: (1) a straightforward syntax for programming productivity, (2) usable data structures for data-oriented programming (DOP) for performance, and (3) a simple interface to the open-source C++ Kokkos library for portability and memory management across CPUs and GPUs. The portability across architectures with a single code implementation is achieved by automatically switching between diverse fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. The MATAR library solves many longstanding challenges associated with easily writing software that can run in parallel on any computer architecture. This work benefits projects seeking to write new C++ codes while also addressing the challenges of quickly making existing Fortran codes performant and portable over modern computer architectures with minimal syntactical changes from Fortran to C++. We demonstrate the feasibility of readily writing new C++ codes and modernizing existing codes with MATAR to be performant, parallel, and portable across diverse computer architectures.

97 MATHEMATICS AND COMPUTING↗

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗