Search NASA⌕ Search

SEARCH · Search NASA

Results for “IRIS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Test cavity and Iris-to-Coax transition for tuning and high-power verification of SNS DTL iris couplers

The Spallation Neutron Source (SNS) Drift Tube Linac (DTL) employs iris couplers to efficiently deliver RF power into the accelerating structure. To support the development, tuning, and high‑power conditioning of these couplers prior to installation in the actual DTLs, a dedicated test cavity and an iris‑to‑coaxial transition structure have been designed. This work presents the electromagnetic design, simulation, and optimization of the test setup, enabling precise characterization of the iris coupler’s performance. The transition structure allows for tuning of the iris opening dimensions without requiring a waveguide taper or full‑size waveguide transitions, while maintaining impedance matching between the coaxial feed and the iris geometry to minimize reflection and power loss. During low‑power tests, the iris opening di-mensions can be evaluated using the iris‑to‑coax transi-tion attached to the test cavity. For high‑power condi-tioning, full‑size waveguides with ceramic vacuum win-dows are connected to the test cavity to replicate opera-tional conditions. Key design parameters were optimized using computer-aided simulation, and sensitivity studies were conducted to assess the impact of mechanical toler-ances on RF performance. The resulting test platform provides a reliable and efficient means for tuning and validating iris couplers, contributing to improved opera-tional stability in the SNS DTL.

Lee, Sung-Woo [ORNL] (ORCID:000000030915835X)↗

IRIS-MEMFLOW: Data Flow-Enabled Portable Memory Orchestration in IRIS Runtime for Diverse Heterogeneity

Task-based programming models and execution paradigms provide a means to decompose a computation by expressing it as a graph in which each node represents a specific computation operating on memory objects and the edges define the dependencies in the execution flow. In this execution model, independent nodes in the graph can be executed concurrently in different computing devices, making it suitable for heterogeneous systems in which computing devices with different architectures coexist. However, careful memory orchestration across heterogeneous devices is needed because copies of the same memory object may reside in multiple devices during execution. Manually ensuring such an orchestration is quite challenging. Not only must an application developer guard against race conditions, but they must also optimize data movement between the host and devices because unnecessary data movement significantly impacts performance. To mitigate these challenges, we enhance the IRIS heterogeneous runtime and introduce IRIS-MEMFLOW–a data flow–enabled portable memory abstraction for seamlessly orchestrating memory in diverse heterogeneous computing environments. By using data-flow analysis, IRIS-MEMFLOW guards against race conditions while multiple heterogeneous devices access memory objects. IRIS-MEMFLOW also optimizes data movement between the host and devices without manual intervention. As a result, IRIS provides improved programming productivity, performance, and portability for multidevice heterogeneous executions in high-performance computing and cloud systems that run diverse architectures from different vendors. The efficacy of IRIS-MEMFLOW is evaluated through experiments that show its capability in terms of programming productivity, multidevice heterogeneity, portability, and low overhead versus the state of the art.

Monil, M. A. H. [ORNL] (ORCID:0000000334194037)↗

Q-IRIS: The Evolution of the IRIS Task-Based Runtime to Enable Classical-Quantum Workflows

Extreme heterogeneity in emerging HPC systems are starting to include quantum accelerators, motivating runtimes that can coordinate between classical and quantum workloads. We present a proof-of-concept hybrid execution framework integrating the IRIS asynchronous task-based runtime with the XACC quantum programming framework via the Quantum Intermediate Representation Execution Engine (QIR-EE). IRIS orchestrates multiple programs written in the quantum intermediate representation (QIR) across heterogeneous backends (including multiple quantum simulators), enabling concurrent execution of classical and quantum tasks. Although not a performance study, we report measurable outcomes through the successful asynchronous scheduling and execution of multiple quantum workloads. To illustrate practical runtime implications, we decompose a four-qubit circuit into smaller subcircuits through a process known as quantum circuit cutting, reducing per-task quantum simulation load and demonstrating how task granularity can improve simulator throughput and reduce queueing behavior -- effects directly relevant to early quantum hardware environments. We conclude by outlining key challenges for scaling hybrid runtimes, including coordinated scheduling, classical-quantum interaction management, and support for diverse backend resources in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

ESnet Requirements Review Program Through the IRI Lens: A Meta-Analysis of Workflow Patterns Across DOE Office of Science Programs (Final Report)

The Department of Energy (DOE) ensures America’s security and prosperity by addressing its energy, environmental, and nuclear challenges through transformative science and technology solutions. The DOE’s Office of Science (SC) delivers groundbreaking scientific discoveries and major scientific tools that transform our understanding of nature and advance the energy, economic, and national security of the United States. The SC’s programs advance DOE mission science across a wide range of disciplines and have developed the research infrastructure needed to remain at the forefront of scientific discovery. The DOE SC’s world-class research infrastructure — exemplified by the 28 SC scientific user facilities — provides the research community with premier observational, experimental, computational, and network capabilities. Each user facility is designed to provide unique capabilities to advance core DOE mission science for its sponsor SC program and to stimulate a rich discovery and innovation ecosystem. Research communities gather and flourish around each user facility, bringing together diverse perspectives. A hallmark of many facilities is the large population of students, postdoctoral researchers, and early-career scientists who contribute as full-fledged users. These facility staff and users collaborate over years to devise new approaches to utilizing the user facility’s core capabilities. The history of the SC user facilities has many examples of wildly inventive researchers challenging operational orthodoxy to pioneer new vistas of discovery; for example, the use of the synchrotron X-ray light sources for study of proteins and other large biological molecules. This continual reinvention of the practice of science — as users and staff forge novel approaches expressed in research workflows — unlocks new discoveries and propels scientific progress. Within this research ecosystem, the high-performance computing (HPC) and networking user facilities stewarded by SC’s Advanced Scientific Computing Research (ASCR) program play a dynamic cross-cutting role, enabling complex workflows demanding high performance data, networking, and computing solutions. The DOE SC’s three HPC user facilities and the Energy Sciences Network (ESnet) high-performance research network serve all of the SC’s programs as well as the global research community. Argonne Leadership Computing Facility (ALCF), the National Energy Research Scientific Computing Center (NERSC), and Oak Ridge Leadership Computing Facility (OLCF) conceive, build, and provide access to a range of supercomputing, advanced computing, and large-scale data-infrastructure platforms, while ESnet interconnects DOE SC research infrastructure and enables seamless exchange of scientific data. All four facilities operate testbeds to expand the frontiers of computing and networking research. Together, the ASCR facilities enterprise seeks to understand and meet the needs and requirements across SC and DOE domain science programs and priority efforts, highlighted by the formal requirements reviews (RRs) methodology. In recent years, the research communities around the SC user facilities have begun experimenting with and demanding solutions integrated with HPC and data infrastructure. This rise of integrated-science approaches is documented in many community and high-level government reports. At the dawn of the era of exascale science and the acceleration of artificial intelligence (AI) innovation, there is a broad need for integrated computational, data, and networking solutions. In response to these drivers, DOE has developed a vision for an Integrated Research Infrastructure (IRI): To empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation.

42 ENGINEERING↗

IRIS-GNN: Leveraging Graph Neural Networks for Scheduling on Truly Heterogeneous Runtime Systems

The diversity of accelerators in computer systems poses significant challenges for software developers, such as managing vendor-specific compiler toolchains, code fragmentation requiring different kernel implementations, and performance portability issues. To address these, the Intelligent Runtime System (IRIS) was developed. IRIS works across various systems, from smartphones to supercomputers, enabling automatic performance scaling based on available accelerators. It introduces abstract tasks for seamless execution transitions between accelerators while ensuring memory consistency and task dependencies. Although IRIS simplifies system details, optimal dynamic scheduling still requires user input to understand workload structures. To address this, we introduce a new scheduling policy for IRIS, termed IRIS-GNN, which is the first IRIS hybrid policy that operates in conjunction with the dynamic policies. This policy employs a Graph-Neural Network (GNN) to conduct Graph Classification of any task graphs submitted to IRIS. This GNN analyzes the structure and attributes of the task graph, categorizing it as either locality, concurrency, or mixed. This classification subsequently guides the selection of the dynamic policy used by IRIS. We provide a comparison of the performance of IRIS-GNN against the complete spectrum of IRIS’s dynamic policies, assess the overhead introduced by the GNN within this scheduling framework, and ultimately explore its practical application in real-world scenarios.

Johnston, Beau↗

IRIS: Exploring Performance Scaling of the Intelligent Runtime System and its Dynamic Scheduling Policies

High-Performance Computing is becoming increasingly heterogeneous, relying on a diverse mix of hardware to achieve good performance. Paradoxically, current drivers and frameworks for these devices typically require separate languages and implementations for each vendor. Furthermore, there are few tools and little support to schedule codes between these devices in a truly heterogeneous manner-partly because of this fragmentation between vendors and the languages each supports. To overcome both limitations, the Intelligent Runtime System (IRIS) was developed. It allows a common task abstraction to automatically be shared among contemporary vendors and is run from a single host-side API. At runtime, IRIS queries the host system and registers which frameworks and drivers are available, these determine which kernels can be used by the scheduler-CPUs via OpenMP, Nvidia GPUs (CUDA), AMD GPUs (HIP), and Intel and Xilinx FPGAs with OpenCL. IRIS enables tasks to be scheduled to any heterogeneous device and resolves to the appropriate kernel binary at runtimeit only uses the devices supported by the system on which it is run. IRIS supports single-task and graph-based expressions of dependencies of tasks. Additionally, IRIS features a range of dynamic scheduling policies, allowing complex chains of tasks and interactions to be executed, relieving the programmer/user from considering the system to assign tasks to devices optimally. This paper presents the peak performance attainable by IRIS over a range of systems-each with different numbers and types of accelerator devices, it highlights the flexibility of IRIS since these devices are truly heterogeneous, relying on different backends (drivers, frameworks, and languages) which historically required unique implementations to utilize them. We then use this peak performance as a baseline to compare increasingly complex chains of tasks (with increasingly complex task dependencies) and evaluate how IRIS copes. Finally, we consider the performance of different IRIS scheduling policies on this range of task graphs.

Johnston, Beau↗

2025 ASMS Investigation of the Collision-Induced Dissociation Mechanism of Protonated TODGA with IRIS

Title (20 words): Investigation of the Collision-Induced Dissociation Mechanism of Protonated TODGA with IRIS Introduction (120 words): One of the challenges facing wide-spread adoption of nuclear power is the development of efficient separation processes for used nuclear fuel. The molecules in separation processes are subjected to an extreme environment due to the high radiation fields from the used fuel and highly acidic media used for fuel dissolution, which results in significant molecular degradation, leading to reduced process efficiency. These degradation products must be identified and studied so mitigation strategies can be developed to maintain process efficiency. However, complex systems can have many degradation products, complicating identification. Untargeted analysis tools could be used to understand radiation chemistry in complex systems. However, this would necessitate improved understanding of the gas-phase fragmentation mechanisms of fuel cycle molecules like tetraoctyldiglycolamide (TODGA). Methods (120 words): The gas-phase fragmentation of protonated TODGA was investigated using collision-induced dissociation (CID), resonance ejection, and infrared ion spectroscopy (IRIS). CID and resonance ejection experiments were conducted using a Bruker Daltonics (Bremen, Gemany) SolariX XR fourier transform ion cyclotron resonance (FT-ICR) mass spectrometer. IRIS spectra of protonated TODGA and its two CID fragmentation products were measured using a modified Bruker amaZon Speed ETD 3D quadrupole ion trap mass spectrometer coupled to the Free Electron Lasers for Infrared eXperiments (FELIX) free electron laser. Measured spectra were compared with density functional theory (DFT) calculations using the Gaussian 16, Revision C.02 software package with the ?B97X-D functional and def2-TZVPP basis sets. Candidate structures were generated using the CREST 3.0 conformational sampling software tool. Preliminary Data (300 words): Collision-induced dissociation of protonated TODGA ([C36H73N2O3]+, m/z=581.562) results two fragment ions, one at m/z=340.285 assigned as [C20H38NO3]+ and the other at m/z=312.290, assigned as [C19H38NO2]+. Based on the assigned formula and the structure of protonated TODGA, the fragment at m/z=340.285 is likely formed from elimination of neutral dioctylamine. Comparison of the IRIS spectrum of m/z=340.285 with DFT predictions suggests it contains a ring structure, and is assigned as N-octyl-N-(6-oxo-1,4-dioxan-2-ylidene)octan-1-aminium. Based on this structure and the structure of protonated TODGA, we hypothesize this fragment formed from elimination of neutral dioctylamine followed by a ring closure mechanism. Comparison of the IRIS spectrum of the fragment at m/z=312.290 with DFT predictions also indicated the presence of a ring structure, assigned as N-(1,3-dioxolan-4-ylidene)-N-octyloctan-1-aminium. This product could be formed from elimination of carbon monoxide from the ring of m/z=340.285 as a sequential fragmentation or formed directly from protonated TODGA via elimination of neutral N,N-dioctylformamide followed by a ring closure. Resonance ejection experiments where m/z=340 was continuously ejected from the IRC cell showed no decrease in intensity of m/z=312.290 across several collision energies, suggesting that the later, direct formation mechanism, dominates. The location of the ionizing proton in protonated TODGA is important for modeling the fragmentation mechanisms. DFT calculations suggested that the position of bands involving the coupled vibrations of the amide C—N and C=O bonds in TODGA are the most sensitive to proton location. Evaluation of the IRIS spectrum of protonated TODGA suggests that the ionizing proton is located between the two amid oxygens. This protonation location was calculated to lie approximately 30 kJ/mol lower in energy than the next lowest energy location, with the proton located solely on one of the amide oxygens. Novel aspect (20 words): Infrared ion spectroscopy combined with resonance ejection experiments and density functional theory to probe the collision-induced dissociation mechanism of tetraoctyldiglycolamide.

37 - INORGANIC, ORGANIC, PHYSICAL AND ANALYTICAL C↗

IRIS-BLAS: Towards a Performance Portable and Heterogeneous BLAS Library

This paper presents IRIS-BLAS, a novel heterogeneous and performance portable BLAS library. IRIS-BLAS is built on top of the IRIS runtime and multiple vendor and open-source BLAS libraries. It can transparently use all the architectures/devices available in a heterogeneous system, using the appropriate BLAS library based on the task mapping at run time. Thus, IRIS-BLAS is portable across a broad spectrum of architectures and BLAS libraries, alleviating the worry of application developers about modifying the application source code. Even though the emphasis is on portability, IRIS-BLAS provides competitive or even better performance than other state-of-the-art references. Moreover, IRIS-BLAS offers new features such as efficiently using extremely heterogeneous systems composed of multiple GPUs from different hardware vendors.

Miniskar, Narasinga Rao↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

FFTX-IRIS: Towards Performance Portability and Heterogeneity for SPIRAL Generated Code

FFTX-IRIS is a dynamic system to efficiently utilize novel heterogeneous platforms. This system links two next-generation frameworks, FFTX and IRIS, to navigate the complexity of different hardware architectures. FFTX provides a runtime code generation framework for high-performance Fast Fourier Transform kernels. IRIS runtime provides portability and multi-device heterogeneity, allowing computation on any available compute resource. Together, FFTX-IRIS enables code generation, seamless portability, and performance without user involvement. We show the design of the FFTX-IRIS system along with an evaluation of various small FFT benchmarks. We also demonstrate multi-device heterogeneity of FFTX-IRIS with a larger stencil application.

Rao, Sanil↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

A Performance-Portable MultiGPU Implementation of 3D Euler Equations using ProtoX and IRIS

Computational scientists often face challenges when developing and optimizing code for high-performance computing (HPC), especially when trying to leverage GPUs. Given the heterogeneity of the nodes that comprise many modern HPC facilities, considerable demand exists for performance portable solutions for the core computational kernels used in many scientific computing libraries. In this work, we demonstrate a fourth-order finite volume method–based implementation of the Euler equations, which are an integral part of computational fluid dynamics. Our performance-portable multiGPU implementation for Euler equations uses ProtoX to generate kernels and IRIS for portability. ProtoX is a domain-specific language that uses a structured-grid partial differential equation library called Proto as its front end and the SPIRAL code generation system as its back end to generate optimized kernels for different architectures. Optimized kernels generated by ProtoX are orchestrated through the IRIS intelligent runtime system to provide portability. Two levels of optimizations within the IRIS runtime— directed acyclic graph fusion and task fusion—are explored to efficiently utilize computing resources in a multiGPU environment. Performance improvement through these optimizations is showcased by comparing the base ProtoX-IRIS implementation on AMD GPUs (Frontier node) and on NVIDIA GPUs (NVIDIA DGX-1).

Mankad, Het↗

Federated IRI Science Testbed (FIRST): A Concept Note

The Department of Energy’s (DOE’s) vision for an Integrated Research Infrastructure (IRI) is to empower researchers to smoothly and securely meld the DOE’s world-class user facilities and research infrastructure in novel ways in order to radically accelerate discovery and innovation. Performant IRI arises through the continuous interoperability of research workflows with compute, storage, and networking infrastructure, fulfilling researchers’ quests to gain insight from observational and experimental data. Decades of successful research, pilot projects, and demonstrations point to the extraordinary promise of IRI but also indicate the intertwined technological, policy, and sociological hurdles it presents. Creating, developing, and stewarding the conditions for seamless interoperability of DOE research infrastructure, with clear value propositions to stakeholders to opt into an IRI ecosystem, will be the next big step. Governance, funding, and resource allocation are beyond the scope of this document: it seeks to provide a high-level view of potential benefits, focus areas, and the working groups whose formation would further define the testbed’s design, activities, and goals.

97 MATHEMATICS AND COMPUTING↗

Design and fabrication of the waveguide Iris couplers for the Spallation Neutron Source drift tube linac

The Spallation Neutron Source (SNS) employs six cavities in the Drift Tube Linac (DTL) section to accelerate the H- ion beam to 87MeV. Each cavity is energized by a 2.5MW peak power klystron at 402.5MHz using rapid tapered waveguide iris couplers. All six original iris couplers have been in operation without replacement for over two decades. The increased RF power demands of the Proton Power Upgrade (PPU) project and operational problems, including arcing, temperature excursions, and vacuum bursts, have prompted the development of new iris coupler spares. The original iris couplers were made of GlidCop material, which is known to be mechanically strong and thermally stable, but is porous, expensive, and difficult to use in fabrication. To overcome these problems, the new spare couplers use Oxygen-Free Copper (OFC) and stainless steel (SS). This paper will discuss the mechanical, thermal and RF design, as well as challenges in the final coupler fabrication.

Lee, Sung-Woo↗

Traceability of Surface Longwave Irradiance Measurements to SI Using the IRIS Radiometers: Preprint

The World Radiation Center at PMOD/WRC is operated on behalf of the WMO. The Infrared Radiometry Section of the WRC (WRC-IRS) provides traceability of downwelling atmospheric longwave radiation measured with pyrgeometers by comparison to the World Infrared Standard Group (WISG) [1]. As has been discussed previously [2], the current implementation of the WISG measures lower longwave irradiances than the two candidate reference radiometers IRIS (Infrared Integrating Sphere Radiometer) and the ACP (Absolute Cavity Pyrgeometer). To validate the findings reported in [2], additional measurements since that publication in 2014 have been performed and are discussed in this paper. Specifically, Measurements during cloud-free nights at PMOD/WRC between the WISG and IRIS radiometers, A field campaign at the Atmospheric Radiation Monitoring Site (ARM) at Southern Great Plains, Oklahoma with IRIS, ACP, AERI (Atmospheric Emittance Radiance Interferometer) and WISG traceable pyrgeometers, Laboratory comparison between the reference blackbody of PMOD/WRC with the hemispherical blackbody developed by PTB.

ACR↗

IRIS Reimagined: Advancements in Intelligent Runtime System for Task-Based Programming

Task-based programming models are gaining traction in scientific computing. IRIS is a portable runtime system that exploits multiple heterogeneous programming systems and can discover available resources and manage multiple diverse programming systems (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, and OpenMP) simultaneously. It accounts for the constraints of task dependencies and provides customizable scheduling policies to map those tasks to heterogeneous devices. In this paper, we present new capabilities added to IRIS to improve its portability for heterogeneous programming, build-friendliness, and performance efficiency. The new additions include vendor-specific kernel support, a runtime system with a foreign function interface to eliminate writing wrapper or boilerplate code for heterogeneous kernels, an easy-to-use and configurable CMake-based build environment, automatic and efficient data transfers and orchestration, and the Hunter and DAGGER toolchains to evaluate IRIS’s task scheduling algorithms.

Miniskar, Narasinga Rao↗

IRIS: A Portable Runtime System Exploiting Multiple Heterogeneous Programming Systems

Across embedded, mobile, enterprise, and HPC systems, computer architectures are becoming more heterogeneous and complex. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to specialize for each architecture. As we show, all of these approaches critically depend on their runtime system for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive runtime system is essential to increase performance portability and improve user productivity. In this regard, we have designed and implemented IRIS: a portable runtime system exploiting multiple heterogeneous programming systems. IRIS can discover available resources, manage multiple diverse programming systems (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

Kim, Jungwon↗

LaRIS: Targeting Portability and Productivity for LAPACK Codes on Extreme Heterogeneous Systems by Using IRIS

In keeping with the trend of heterogeneity in high-performance computing, hardware manufacturers and vendors are developing new architectures and associated software stacks (e.g., libraries) to harness the best possible performance from commonly used kernels (e.g., linear algebra kernels). However, kernels tuned for one architecture are not portable to others. Moreover, the coexistence of different architectures in a single node makes orchestration difficult. To address these challenges, we introduce LaRIS, a portable framework for LAPACK functionalities. LaRIS ensures a separation between linear algebra algorithms and vendor-library kernels by using the IRIS run time and IRIS-BLAS library. Such abstraction at the algorithm level makes the implementation completely agnostic to the vendor library and architecture. LaRIS uses the IRIS run time to dynamically select the vendor-library kernel and suitable processor architecture at run time. Through LU factorization, we demonstrate that LaRIS can fully utilize different heterogeneous systems by launching and orchestrating different vendor-library kernels without any change in the source code.

Monil, M. A. H.↗