Search NASA⌕ Search

SEARCH · Search NASA

Results for “compiler optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A survey of compiler optimization techniques

Major optimization techniques of compilers are described and grouped into three categories: machine dependent, architecture dependent, and architecture independent. Machine-dependent optimizations tend to be local and are performed upon short spans of generated code by using particular properties of an instruction set to reduce the time or space required by a program. Architecture-dependent optimizations are global and are performed while generating code. These optimizations consider the structure of a computer, but not its detailed instruction set. Architecture independent optimizations are also global but are based on analysis of the program flow graph and the dependencies among statements of source program. A conceptual review of a universal optimizer that performs architecture-independent optimizations at source-code level is also presented.

Schneck, P. B.↗

Retargetable Optimizing Compilers for Quantum Accelerators via a Multi-Level Intermediate Representation

In this work, we present a multi-level quantum-classical intermediate representation (IR) that enables an optimizing, retargetable compiler for available quantum languages. Our work builds upon the Multi-level Intermediate Representation (MLIR) framework and leverages its unique progressive lowering capabilities to map quantum languages to the LLVM machine-level IR. We provide both quantum and classical optimizations via the MLIR pattern rewriting sub-system and standard LLVM optimization passes, and demonstrate the programmability, compilation, and execution of our approach via standard benchmarks and test cases. In comparison to other standalone language and compiler efforts available today, our work results in compile times that are 1000x faster than standard Pythonic approaches, and 5-10x faster than comparative standalone quantum language compilers. Our compiler provides quantum resource optimizations via standard programming patterns that result in a 10x reduction in entangling operations, a common source of program noise. We see this work as a vehicle for rapid quantum compiler prototyping.

43 PARTICLE ACCELERATORS↗

Ciel

Compiler optimizations can alter the numerical results of scientific computing applications. When numerical results differ significantly between compilers, optimization levels, and floating-point hardware, these numerical inconsistencies can impact programming productivity. Ciel is a framework that helps programmers identify locations in the source code that are affected by compiler optimizations in CPU and GPU code. Ciel uses a floating-point precision enhancement strategy, guided by a recursive bisection search algorithm with increasing search granularity, to identify the program expressions that induce numerical inconsistencies due to compiler optimizations.

Miao, Wenjun↗

Tough Errors Are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This report summarizes research performed under the Tough Errors Are no Match (TEAM) project. The primary focus of TEAM has been to research and develop a compilation toolbox leveraging techniques from quantum characterization and control, probabilistic programming, and approximate computing. Our goal was to develop robust protocols that can be integrated into quantum compilers to optimize and enhance the robustness of noisy computation. Here, we provide a summary of TEAM work focused on characterization and control of quantum systems.

97 MATHEMATICS AND COMPUTING↗

Noise-aware circuit compilations for a continuously parameterized two-qubit gateset

State-of-the-art noisy-intermediate-scale quantum processors are currently implemented across a variety of hardware platforms, each with their own distinct gatesets. As such, circuit compilation should not only be aware of but also deeply connect to the native gateset and noise properties of each. Trapped-ion processors are one such platform that provides a gateset that can be continuously parameterized across both one- and two-qubit gates. Here we use the Quantum Scientific Computing Open User Testbed to study noise-aware compilations focused on continuously parameterized two-qubit 𝑍⁢𝑍 gates (based on the Mølmer-Sørensen interaction) using $\scriptsize{SUPERSTAQ}$, a quantum software platform for hardware-aware circuit compiler optimizations. We discuss the realization of 𝑍⁢𝑍 gates with arbitrary angle on the all-to-all connected trapped-ion system. Then we discuss a variety of different compiler optimizations that innately target these 𝑍⁢𝑍 gates and their noise properties. These optimizations include moving from a restricted maximally entangling gateset to a continuously parameterized one, swap mirroring to further reduce the total entangling angle of the operations, focusing the heaviest 𝑍⁢𝑍 angle participation on the best-performing gate pairs, and circuit approximation to remove the least impactful 𝑍⁢𝑍 gates. We demonstrate these compilation approaches on the hardware with randomized quantum volume circuits, observing the potential to realize a larger quantum volume as a result of these optimizations. Using differing yet complementary analysis techniques, we observe the distinct improvements in system performance provided by these noise-aware compilations and study the role of stochastic and coherent error channels for each compilation choice.

Noise↗

QuComm: Optimizing Collective Communication for Distributed Quantum Computing

Distributed quantum computing (DQC) is a scalable way to build a large-scale quantum computing system while the error-prone nonlocal communication between DQC nodes may heavily degrade the fidelity of the distributed quantum program and thus demands specific compiler optimizations. Previous compilers on DQC communication optimization either assumes unlimited communication resource or a few communication qubits due to the hardware limitation. The former compilers may not be efficient when interfacing with communication-resource-constrained DQC hardware while the latter compilers lose the opportunities of optimizing collective communication and routing concurrent communication as they unnecessarily couple limited communication qubits with the implementation of expensive inter-node operations. In this paper, we invent the communication buffer, a communication facility consisting of idle qubits in each compute node, to decouple the execution of inter-node quantum operations from communication qubits: communication qubits are devoted to generating inter-node entanglement while internode operations are conducted in the communication buffer. The communication buffer provides an intermediate layer for inter-node communication and paves the way for collective communication optimization. We then propose QuComm, a buffer-based compiler framework that first performs smart buffer allocation according to communication characteristics of the distributed quantum program and then optimizes and collectively routes inter-node quantum operations. Experimental results on a hierarchical DQC system show that the proposed QuComm can reduce the most expensive inter-node communication request and the latency of various distributed quantum programs by 50.4% and 47.6% on average, respectively.

Wu, Anbang↗

COMPOFF: A Compiler Cost model using Machine Learning to predict the Cost of OpenMP Offloading

The HPC industry is inexorably moving towards an era of extremely heterogeneous architectures, with more devices configured on any given HPC platform and potentially more kinds of devices, some of them highly specialized. Writing a separate code suitable for each target system for a given HPC application is not practical. The better solution is to use directive-based parallel programming models such as OpenMP. OpenMP provides a number of options for offloading a piece of code to devices like GPUs. To select the best option from such options during compilation, most modern compilers use analytical models to estimate the cost of executing the original code and the different offloading code variants. Building such an analytical model for compilers is a difficult task that necessitates a lot of effort on the part of a compiler engineer. Recently, machine learning techniques have been successfully applied to build cost models for a variety of compiler optimization problems. In this paper, we present COMPOFF, a cost model which uses the multi-layer perceptrons to statically estimates the Cost of OpenMP OFFloading. We used six different transformations on a parallel code of Wilson Dslash Operator to support GPU offloading, and we predicted their cost of execution on different GPUs using COMPOFF during compile time. Our results show that this model can predict offloading costs with a root mean squared error in prediction of less than 0.5 seconds. Our preliminary findings indicate that this work will make it much easier and faster for scientists and compiler developers to port legacy HPC applications that use OpenMP to new heterogeneous computing environment.

97 MATHEMATICS AND COMPUTING↗

QGLab v0.0.1

A software program for experimenting with holographic teleportation protocol on quantum computers. The code implements the protocol that was proposed in https://arxiv.org/abs/1911.06314. The code is an end-to-end software solution that facilitates conducting the holographic teleportation experiments on state-of-the-art and emergent generations of QPUs supported by the Qiskit and tket SDKs. The code bundles all stages of an experiment as a single configurable workflow allowing faster development and experimentation cycles. Features: 1. Easy switching between Qiskit and tket quantum compilers. 2. Semi-automatic facilities for finding optimal compilation solutions beyond what Qiskit and tket provide by default. 3. Experiment resolution scaling (based on automatic jobs' batching). 4. Automatic experiment scaling over qubits. 5. Automatic readout error mitigation. 6. Automatic reproducibility analysis. 7. Standalone error-mitigation tools (randomized compiling, mitigation with estimation circuits, zero-noise extrapolation)

Shapoval, Illya↗

Reoptimization of Quantum Circuits via Hierarchical Synthesis

The current phase of quantum computing is in the Noisy Intermediate-Scale Quantum (NISQ) era. On NISQ devices, two-qubit gates such as CNOTs are much noisier than single-qubit gates, so it is essential to minimize their count. Quantum circuit synthesis is a process of decomposing an arbitrary unitary into a sequence of quantum gates, and can be used as an optimization tool to produce shorter circuits to improve overall circuit fidelity. However, the time-to-solution of synthesis grows exponentially with the number of qubits. As a result, synthesis is intractable for circuits on a large qubit scale. In this paper, we propose a hierarchical, block-by-block opti-mization framework, QGo, for quantum circuit optimization. Our approach allows an exponential cost optimization to scale to large circuits. QGo uses a combination of partitioning and synthesis: 1) partition the circuit into a sequence of independent circuit blocks; 2) re-generate and optimize each block using quantum synthesis; and 3) re-compose the final circuit by stitching all the blocks together. We perform our analysis and show the fidelity improvements in three different regimes: small-size circuits on real devices, medium-size circuits on noisy simulations, and large-size circuits on analytical models. Our technique can be applied after existing optimizations to achieve higher circuit fidelity. Further, using a set of NISQ benchmarks, we show that QGo can reduce the number of CNOT gates by 29.9% on average and up to 50% when compared with industrial compiler optimizations such as t|ket). When executed on the IBM Athens system, shorter depth leads to higher circuit fidelity. We also demonstrate the scalability of our QGo technique to optimize circuits of 60+ qubits, Our technique is the first demonstration of successfully employing and scaling synthesis in the compilation tool chain for large circuits. Overall, our approach is robust for direct incorporation in production compiler toolchains to further improve the circuit fidelity.

97 MATHEMATICS AND COMPUTING↗

Proof-Carrying Code with Correct Compilers

In the late 1990s, proof-carrying code was able to produce machine-checkable safety proofs for machine-language programs even though (1) it was impractical to prove correctness properties of source programs and (2) it was impractical to prove correctness of compilers. But now it is practical to prove some correctness properties of source programs, and it is practical to prove correctness of optimizing compilers. We can produce more expressive proof-carrying code, that can guarantee correctness properties for machine code and not just safety. We will construct program logics for source languages, prove them sound w.r.t. the operational semantics of the input language for a proved-correct compiler, and then use these logics as a basis for proving the soundness of static analyses.

Appel, Andrew W.↗

Explicit time integration of finite element models on a vectorized, concurrent computer with shared memory

The implementation of a nonlinear explicit program on a vectorized, concurrent computer with shared memory is described and studied. The conflict between vectorization and concurrency is described and some guidelines are given for optimal block sizes. Several example problems are summarized to illustrate the types of speed-ups which can be achieved by reprogramming as compared to compiler optimization.

Gilbertsen, Noreen D.↗

Software Deployment Process at NERSC: Deploying the Extreme-scale Scientific Software Stack (E4S) Using Spack at the National Energy Research Scientific Computing Center (NERSC)

One of the many benefits of using a high-performance computing (HPC) system at a Department of Energy (DOE) Office of Science (SC) HPC facility is the large number of software products, built and optimized for the system. The HPC center staff and HPC vendors provide optimized software such as libraries and even full scientific applications, ready to be used by users as building blocks to accelerate scientific discovery. Behind each provided packaged software module are a large number of decisions - which compiler, optimizations, variants/options - to build the software on the target system. And, even before the software gets deployed, the software must be developed, tested, and maintained, including deprecating old versions and ensuring compatibility across versions. The software lifecycle is complex and is further convoluted by a web of interdependencies on other software.

97 MATHEMATICS AND COMPUTING↗

Porting numerical integration codes from CUDA to oneAPI: a case study

We present our experience in porting optimized CUDA implementations to oneAPI. We focus on the use case of numerical integration, particularly the CUDA implementations of PAGANI and $m$-Cubes. We faced several challenges that caused performance degradation in the oneAPI ports. These include differences in utilized registers per thread, compiler optimizations, and mappings of CUDA library calls to oneAPI equivalents. After addressing those challenges, we tested both the PAGANI and m-Cubes integrators on numerous integrands of various characteristics. To evaluate the quality of the ports, we collected performance metrics of the CUDA and oneAPI implementations on the Nvidia V100 GPU. We found that the oneAPI ports often achieve comparable performance to the CUDA versions, and that they are at most 10% slower.

97 MATHEMATICS AND COMPUTING↗

QASMTrans: A QASM Quantum Transpiler Framework for NISQ Devices

In quantum computing, transpilation plays a crucial role in converting high-level, machine-independent quantum circuits into circuits specially for a quantum device, considering factors such as basis gate set, topology, error profile, etc. Yet, the efficiency of transpilation remains a significant bottleneck, particularly when dealing with very large QASM level input files. In this paper, we present QASMTrans, a C++ based high-performance quantum transpiler framework that can demonstrate on average 50-100× speedups compared to the internal transpiler of Qiskit. Particularly, for large dense circuits such as ’uccsd n24’ and ’qft n320’ incorporating millions of gates, QASMTrans can successfully transpile in 69s and 31s, respectively, while Qiskit failed to finish in one hour. Using QASMTrans as the baseline, it becomes more feasible to explore much larger design space and impose more comprehensive compiler optimizations.

Hua, Fei↗

Automatic inspection of program state in an uncooperative environment

Abstract The program state is formed by the values that the program manipulates. These values are stored in the stack, in the heap, or in static memory. The ability to inspect the program state is useful as a debugging or as a verification aid. Yet, there exists no general technique to insert inspection points in type‐unsafe languages such as C or C++. The difficulty comes from the need to traverse the memory graph in a so‐called uncooperative environment. In this article, we propose an automatic technique to deal with this problem. We introduce a static code transformation approach that inserts in a program the instrumentation necessary to report its internal state. Our technique has been implemented in LLVM. It is possible to adjust the granularity of inspection points trading precision for performance. In this article, we demonstrate how to use inspection points to debug compiler optimizations; to augment benchmarks with verification code; and to visualize data structures.

Magalhães, José Wesley de Souza↗

Boosting RDataFrame performance with transparent bulk event processing

RDataFrame is ROOT’s high-level interface for Python and C++ data analysis. Since it first became available, RDataFrame adoption has grown steadily and it is now poised to be a major component of analysis software pipelines for LHC Run 3 and beyond. Thanks to its design inspired by declarative programming principles, RDataFrame enables the development of highperformance, highly parallel analyses without requiring expert knowledge of multi-threading and I/O: user logic is expressed in terms of self-contained, small computation kernels tied together by a high-level API. This design completely decouples analysis logic from its actual execution, and opens several interesting avenues for workflow optimization. In particular, in this work we explore the benefits of moving internal data processing from an event-by-event to a bulkby-bulk loop. This refactoring dramatically reduces the framework’s runtime overheads; in collaboration with the I/O layer it improves data access patterns; it exposes information that optimizing compilers might use to auto-vectorize the invocation of user-defined computations; finally, while existing user-facing interfaces remain unaffected, it becomes possible to additionally offer interfaces that explicitly expose bulks of events, useful e.g. for the injection of GPU kernels into the analysis workflow. In order to inform similar future R&D, design challenges will be presented, as well as an investigation of the relevant timememory trade-off backed by novel performance benchmarks.

Guiraud, Enrico↗