Search NASA⌕ Search

SEARCH · Search NASA

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

On the practical usefulness of the Hardware Efficient Ansatz

Variational Quantum Algorithms (VQAs) and Quantum Machine Learning (QML) models train a parametrized quantum circuit to solve a given learning task. The success of these algorithms greatly hinges on appropriately choosing an ansatz for the quantum circuit. Perhaps one of the most famous ansatzes is the one-dimensional layered Hardware Efficient Ansatz (HEA), which seeks to minimize the effect of hardware noise by using native gates and connectives. The use of this HEA has generated a certain ambivalence arising from the fact that while it suffers from barren plateaus at long depths, it can also avoid them at shallow ones. In this work, we attempt to determine whether one should, or should not, use a HEA. We rigorously identify scenarios where shallow HEAs should likely be avoided (e.g., VQA or QML tasks with data satisfying a volume law of entanglement). More importantly, we identify a Goldilocks scenario where shallow HEAs could achieve a quantum speedup: QML tasks with data satisfying an area law of entanglement. We provide examples for such scenario (such as Gaussian diagonal ensemble random Hamiltonian discrimination), and we show that in these cases a shallow HEA is always trainable and that there exists an anti-concentration of loss function values. Our work highlights the crucial role that input states play in the trainability of a parametrized quantum circuit, a phenomenon that is verified in our numerics.

97 MATHEMATICS AND COMPUTING↗

Hardware-Efficient Monitoring of I/O Signals

In this invention, command and monitor functionality is moved between the two independent pieces of hardware, in which one had been dedicated to command and the other had been dedicated to monitor, such that some command and some monitor functionality appears in each. The only constraint is that the monitor for signal cannot be in the same hardware as the command I/O it is monitoring. The splitting of the command outputs between independent pieces of hardware may require some communication between them, i.e. an intra-switch trunk line. This innovation reduces the amount of wasted hardware and allows the two independent pieces of hardware to be designed identically in order to save development costs.

Driscoll, Kevin R.↗

Hardware-Efficient Quantum Optimization Layered Algorithms and Experiments

Quantum optimization algorithms, such as QAOA, that implement parametrized stochastic optimization solvers attempt to identify low-energy solutions of Ising systems by exploiting available quantum effects in noisy-intermediate scale machines. Engineering a well-performing parametrized quantum optimization circuit is indeed an exercise in balancing the trade-off between expressivity and implementation complexity. We show that, for MaxCut QAOA circuits defined on native hardware topology (Rigetti’s Aspen Quantum Processors), error-mitigation techniques recover simulated features of the noiseless theory. Moreover, we explore a design space for QAOA-like ansatze that perform well in theory as well as in hardware for fully-connected problems. We also discuss how efficient coherence and entanglement detection methods that could be coupled with quantum optimization experiments require only linear overhead in benchmarking time.

quantum computing↗

Hardware efficient monitoring of input/output signals

A communication device comprises first and second circuits to implement a plurality of ports via which the communicative device is operable to communicate over a plurality of communication channels. For each of the plurality of ports, the communication device comprises: command hardware that includes a first transmitter to transmit data over a respective one of the plurality of channels and a first receiver to receive data from the respective one of the plurality of channels; and monitor hardware that includes a second receiver coupled to the first transmitter and a third receiver coupled to the respective one of the plurality of channels. The first circuit comprises the command hardware for a first subset of the plurality of ports. The second circuit comprises the monitor hardware for the first subset of the plurality of ports and the command hardware for a second subset of the plurality of ports.

Driscoll, Kevin R.↗

Hardware-Efficient Quantum Phase Estimation via Local Control

Quantum phase estimation plays a central role in quantum simulation as it enables the study of spectral properties of many-body quantum systems. Most variants of the phase estimation algorithm require the application of the global unitary evolution conditioned on the state of one or more auxiliary qubits, posing a significant challenge for current quantum devices. In this work, we present an approach to quantum phase estimation that uses only locally controlled operations, resulting in a significantly reduced circuit depth. At the heart of our approach are efficient routines to measure the complex phase of the expectation value of the time-evolution operator, the so-called Loschmidt echo, for both circuit dynamics and Hamiltonian dynamics. By tracking changes in the phase during the dynamics, the routines trade circuit depth for increased sampling cost and classical postprocessing. Our approach does not rely on reference states and is applicable to any efficiently preparable state, regardless of its correlations. We provide a comprehensive analysis of the sample complexity and illustrate the results with numerical simulations. Our methods offer a practical pathway for measuring spectral properties in large many-body quantum systems using current quantum devices.

Schiffer, Benjamin F. [Max Planck Institute of Qua↗

Cascade Error Projection: An Efficient Hardware Learning Algorithm

A new learning algorithm termed cascade error projection (CEP) is presented. CEP is an adaption of a constructive architecture from cascade correlation and the dynamical stepsize of A/D conversion from the cascade back propagation algorithm.

learning algorithm pattern recognition cascade err↗

Fault-Tolerant Operation of Bosonic Qubits with Discrete-Variable Ancillae

Fault-tolerant quantum computation with bosonic qubits often necessitates the use of noisy discrete-variable ancillae. In this work, we establish a comprehensive and practical fault-tolerance framework for such a hybrid system and synthesize it with fault-tolerant protocols by combining bosonic quantum error correction (QEC) and advanced quantum control techniques. We introduce essential building blocks of error-corrected gadgets by leveraging ancilla-assisted bosonic operations using a generalized variant of path-independent quantum control. Using these building blocks, we construct a universal set of error-corrected gadgets that tolerate a single-photon loss and an arbitrary ancilla fault for four-legged cat qubits. Notably, our construction requires only dispersive coupling between bosonic modes and ancillae, as well as beam-splitter coupling between bosonic modes, both of which have been experimentally demonstrated with strong strengths and high accuracy. Moreover, each error-corrected bosonic qubit is comprised of only a single bosonic mode and a three-level ancilla, featuring the hardware efficiency of bosonic QEC in the full fault-tolerant setting. We numerically demonstrate the feasibility of our schemes using current experimental parameters in the circuit-QED platform. Finally, we present a hardware-efficient architecture for fault-tolerant quantum computing by concatenating the four-legged cat qubits with an outer qubit code utilizing only beam-splitter couplings. Our estimates suggest that the overall noise threshold can be reached using existing hardware. These developed fault-tolerant schemes extend beyond their applicability to four-legged cat qubits and can be adapted for other rotation-symmetrical codes, offering a promising avenue toward scalable and robust quantum computation with bosonic qubits. Published by the American Physical Society 2024

Physics↗

Tailor : Altering Skip Connections for Resource-Efficient Inference

Deep neural networks use skip connections to improve training convergence. However, these skip connections are costly in hardware, requiring extra buffers and increasing on- and off-chip memory utilization and bandwidth requirements. In this article, we show that skip connections can be optimized for hardware when tackled with a hardware-software codesign approach. We argue that while a network’s skip connections are needed for the network to learn, they can later be removed or shortened to provide a more hardware-efficient implementation with minimal to no accuracy loss. We introduceTailor, a codesign tool whose hardware-aware training algorithm gradually removes or shortens a fully trained network’s skip connections to lower the hardware cost.Tailorimproves resource utilization by up to 34% for block random access memories (BRAMs), 13% for flip-flops (FFs), and 16% for look-up tables (LUTs) for on-chip, dataflow-style architectures.Tailorincreases performance by 30% and reduces memory bandwidth by 45% for a two-dimensional processing element array architecture.

Computer Science↗

Extended Logic Intelligent Processing System for a Sensor Fusion Processor Hardware

The paper presents the hardware implementation and initial tests from a low-power, highspeed reconfigurable sensor fusion processor. The Extended Logic Intelligent Processing System (ELIPS) is described, which combines rule-based systems, fuzzy logic, and neural networks to achieve parallel fusion of sensor signals in compact low power VLSI. The development of the ELIPS concept is being done to demonstrate the interceptor functionality which particularly underlines the high speed and low power requirements. The hardware programmability allows the processor to reconfigure into different machines, taking the most efficient hardware implementation during each phase of information processing. Processing speeds of microseconds have been demonstrated using our test hardware.

Stoica, Adrian↗

MIL-M-38510/470 test vectors: Fault detection efficiency measurement via hardware fault simulation

The stuck fault detection efficiency of the test vectors developed for the MIL-M-38510/470 NASA was measured using a hardware stuck fault simulator for the 1802 microprocessor. Thirty-nine stuck faults were not detected out of a total of 874 injected into the combinatorial and sequential parts of the microprocessor. Since undetected faults can create catastrophic errors in equipment designed for high reliability applications, it is recommended that the MIL-M-38510/470 NASA be enhanced with additional test vectors so as to achieve 100% stuck fault detection efficiency.

Timoc, C. C.↗

Neural architecture codesign for fast physics applications

We develop a pipeline to streamline neural architecture codesign for physics applications to reduce the need for ML expertise when designing models for novel tasks. Our method employs neural architecture search and network compression in a two-stage approach to discover hardware efficient models. This approach consists of a global search stage that explores a wide range of architectures while considering hardware constraints, followed by a local search stage that fine-tunes and compresses the most promising candidates. We exceed performance on various tasks and show further speedup through model compression techniques such as quantization-aware-training and neural network pruning. We synthesize the optimal models to high level synthesis code for FPGA deployment with the hls4ml library. Additionally, our hierarchical search space provides greater flexibility in optimization, which can easily extend to other tasks and domains. We demonstrate this with two case studies: Bragg peak finding in materials science and jet classification in high energy physics, achieving models with improved accuracy, smaller latencies, or reduced resource utilization relative to the baseline models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

EMU processing - A myth dispelled

The refurbishment-and-checkout 'processing' activities entailed by the Space Shuttle Extravehicular Mobility Units (EMUs) are currently significantly more modest, at 1050 man-hours, than when Space Shuttle services began (involving about 4000 man-hours). This great improvement in hardware efficiency is due to the design or modification of test rigs for simplification of procedures, as well as those procedures' standardization, in conjunction with an increase in hardware confidence which has allowed the extension of inspection, service, and testing intervals. Recent simplification of the hardware-processing sequence could reduce EMU processing requirements to 600 man-hours in the near future.

Peacock, Paul R.↗

Machine Learning Enabled Position Detection for 6.78 MHz UAV Wireless Power Transfer System

This paper presents a novel supervised machine learning (SML) approach for accurate position detection of the receiver coil in wireless power transfer (WPT) systems using only secondary-side electrical measurements, with applications in autonomous unmanned aerial vehicle (UAV) charging. The proposed method trains a supervised learning model to map measured secondary-side voltage and current features to the receiver’s spatial position with high precision. This enables an autonomous UAV to determine its location relative to the primary coil center, the optimal position for maximizing wireless charging efficiency. The sensing method is fully integrated into a standard WPT system, utilizing the same primary and secondary coils for both power transfer and position detection, thereby eliminating additional sensing hardware. The use of a 6.78 MHz operating frequency enhances positional sensitivity, as high-frequency near-field electromagnetic fields respond strongly to small spatial variations. Experimental validation is performed on a 30 W scaled prototype featuring a 210 mm × 140 mm primary coil, a 50 mm × 80 mm receiver coil, and a 15 mm air gap. Results demonstrate reliable position estimation and a strong correlation between predicted position and optimal coil alignment. This integrated framework unifying position detection and wireless charging offers a promising foundation for future autonomous electric vertical takeoff and landing (eVTOL) systems, enabling compact, hardware-efficient, and high-accuracy charging solutions.

Colak, Kerim [New York University]↗

Status of utility-interactive photovoltaic power conditioning technology

Design options for utility-interactive photovoltaic power conditioning technology for unit ratings from 2kW to 5 MW are compared. Line- and self-commutated inverter designs for both single and three-phase applications are described. Efficiency, weight, and cost projections are provided for comparing the design options. New circuit designs that take advantage of advances in power semiconductor devices are found to be the most promising. Hardware efficiencies from 95 percent for single phase to 98 percent for three-phase applications are found.

Key, T. S.↗

Unifying Combinatorial and Graphical Methods in Artificial Intelligence

Recently, a new graph Laplacian, called the inner product Laplacian, was introduced which generalizes many existing Laplacians, including the normalized and combinatorial Laplacian and their weighted variants. The key observation behind the inner product Laplacian is that by defining appropriate inner product spaces on the vertices and edges, the standard Laplacians can be recovered as Hodge Laplacians over the simplicial complex formed by the edges and vertices. These inner product spaces form a natural way to incorporate non-combinatorial information into the definition of a domain-specific Laplacian. In particular, in contrast to current domain-specific weighting schemes which rely solely on edge weights, information regarding the similarity of non-adjacent vertices and arbitrary pairs of edges can be effectively incorporated into the Laplacian. In order to illustrate this approach we consider the problem of calculating the potential energy of an atomistic configuration using Graph Neural Networks. In comparison with start-of-the-art approaches, such as SchNet, our approach replaces a learned (via auto-encoder) representation of the atom types with an inner product space on atoms based on scientific knowledge (e.g., electronegativity). We will illustrate how this approach captures key chemical properties of the molecules and compare the energy calculations with state-of-the-art neural network approaches. However, to compute the resulting Laplacian involves a mixture of sparse and dense matrix computation and yields a dense matrix as the basis for the graph convolution. This dense convolutional kernel necessitates moving away from the standard message passing framework for graph neural networks and increases the computational cost of applying the kernel. In order to mitigate these costs we investigate means of leveraging the mixed sparse and dense computations to reduce the overall computational cost and how these approaches can be automatically transferred to energy efficient hardware (e.g., field programmable gate arrays (FPGAs)).

97 MATHEMATICS AND COMPUTING↗

Supervised Learning-Based Spatial Position Estimation with Vertical Displacement for Hovering UAV Wireless Power Transfer

This study presents a supervised learning-based spatial position estimation approach for wireless power transfer (WPT) systems supporting hovering unmanned aerial vehicle (UAV) charging. Unlike stationary charging scenarios, hovering UAVs introduce continuous lateral misalignment and vertical displacement, leading to variations in magnetic coupling and reduced power transfer efficiency. To address this challenge, the proposed method estimates the relative spatial position of the receiver coil using only electrical measurements obtained at the secondary side. A supervised learning model is trained to map output voltage and current features to spatial coordinates, enabling position awareness without requiring external sensors, vision systems, or communication links. The sensing functionality is inherently integrated into the WPT system, allowing simultaneous power transfer and localization through the same magnetic interface. Experimental validation is conducted on a laboratory-scale prototype under varying lateral offsets and air-gap conditions. In addition, spline-based interpolation is employed to increase spatial data density for training. The results demonstrate that the proposed framework can capture spatial variations associated with both lateral and vertical displacement, providing reliable position estimation under hovering conditions. This work establishes a hardware-efficient, sensorless solution for UAV wireless charging and serves as a baseline for advanced data-driven position estimation methods in dynamic WPT systems.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

Nonunitary Variational Quantum Eigensolver with the Localized Active Space Method and Cost Mitigation

Accurately describing strongly correlated systems with affordable quantum resources remains a central challenge for quantum chemistry applications on near and intermediate term quantum computers. The localized active space self-consistent field (LASSCF) approximates the complete active space self-consistent field (CASSCF) by generating active space-based wave functions within specific fragments while treating interfragment correlation with mean-field approach, hence is computationally less expensive. Hardware-efficient ansatzes (HEA) offer affordable and shallower circuits, yet they often fail to capture the necessary correlation. Previously, Jastrow-factor-inspired nonunitary qubit operators were proposed to use with HEA for variational quantum eigensolver (VQE) calculations (so-called nuVQE), as they do not increase circuit depths and recover correlation beyond the mean-field level for Hartree–Fock initial states. Here, in this study, we explore running nuVQE with LASSCF as the initial state. The method, named LAS-nuVQE, is shown to recover interfragment correlations, reach chemical accuracy with a small number of gates (<70) in both H 4 and square cyclobutadiene (C 4 H 4 ), and produces more accurate energetics than its HEA counterparts at all circuit depths. To further address the inherent symmetry-breaking in HEA, we implemented spin-constrained LAS-nuVQE to extend the capabilities of HEA further and show spin-pure results for square cyclobutadiene. We also mitigate the increased measurement overhead of nuVQE via Pauli grouping and shot-frugal sampling, reducing measurement costs by up to 2 orders of magnitude compared to ungrouped operator, and show that one can achieve better accuracy with a small number of shots (10 3–4 ) per one expectation value calculation compared to noiseless simulations with one or two orders of magnitude more shots. Finally, wall clock time estimates show that, with our measurement mitigation protocols, nuVQE becomes a cheaper and more accurate alternative than vanilla VQE with HEA. Taken together, these developments illustrate a practical pathway toward performing multireference chemical simulations with accuracy and affordable resources on today’s quantum hardware, achieving both accuracy and affordability in challenging correlated systems.

Wang, Qiaohong [Univ. of Chicago, IL (United State↗

End-to-End Workflow for Machine-Learning-Based Qubit Readout With QICK and hls4ml

In this article, we present an end-to-end workflow for superconducting qubit readout that embeds codesigned neural networks into the quantum instrumentation control kit (QICK). Capitalizing on the custom firmware and software of the QICK platform, which is built on Xilinx radiofrequency system-on-chip field-programmable gate arrays (FPGAs), we aim to leverage machine learning (ML) to address critical challenges in qubit readout accuracy and scalability. The workflow utilizes the hls4ml package and employs quantization-aware training to translate ML models into hardware-efficient FPGA implementations via user-friendly Python application programming interfaces. We experimentally demonstrate the design, optimization, and integration of an ML algorithm for single transmon qubit readout, achieving 96% single-shot fidelity with a latency of 32.25 ns and less than 16% FPGA lookup table resource utilization. Our results offer the community an accessible workflow to advance ML-driven readout and adaptive control in quantum information processing applications.

42 ENGINEERING↗