Search NASA⌕ Search

SEARCH · Search NASA

Results for “communication lower bounds”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Communication Lower Bounds and Optimal Algorithms for Multiple Tensor-Times-Matrix Computation

Multiple tensor-times-matrix (Multi-TTM) is a key computation in algorithms for computing and operating with the Tucker tensor decomposition, which is frequently used in multidimensional data analysis. Here, we establish communication lower bounds that determine how much data movement is required (under mild conditions) to perform the Multi-TTM computation in parallel. The crux of the proof relies on analytically solving a constrained, nonlinear optimization problem. We also present a parallel algorithm to perform this computation that organizes the processors into a logical grid with twice as many modes as the input tensor. We show that, with correct choices of grid dimensions, the communication cost of the algorithm attains the lower bounds and is therefore communication optimal. Finally, we show that our algorithm can significantly reduce communication compared to the straightforward approach of expressing the computation as a sequence of tensor-times-matrix operations when the input and output tensors vary greatly in size.

HBL-inequalities↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

Capacities of Entanglement Distribution From a Central Source

Distribution of entanglement is an essential task in quantum information processing and the realization of quantum networks. In our work, we theoretically investigate the scenario where a central source prepares an N -partite entangled state and transmits each entangled subsystem to one of N receivers through noisy quantum channels. The receivers are then able to perform local operations assisted by unlimited classical communication to distill target entangled states from the noisy channel output. In this operational context, we define the EPR distribution capacity and the GHZ distribution capacity of a quantum channel as the largest rates at which Einstein-Podolsky-Rosen (EPR) states and Greenberger-Horne-Zeilinger (GHZ) states can be faithfully distributed through the channel, respectively. We establish lower and upper bounds on the EPR distribution capacity by connecting it with the task of assisted entanglement distillation. We also construct an explicit protocol consisting of a combination of a quantum communication code and a classical-post-processing-assisted entanglement generation code, which yields a simple achievable lower bound for generic channels. As applications of these results, we give an exact expression for the EPR distribution capacity over two erasure channels and bounds on the EPR distribution capacity over two generalized amplitude damping channels. We also bound the GHZ distribution capacity, which results in an exact characterization of the GHZ distribution capacity when the most noisy channel is a dephasing channel.

42 ENGINEERING↗

Quantum Entanglement between Optical and Microwave Photonic Qubits

Entanglement is an extraordinary feature of quantum mechanics. Sources of entangled optical photons were essential to test the foundations of quantum physics through violations of Bell’s inequalities. More recently, entangled many-body states have been realized via strong nonlinear interactions in microwave circuits with superconducting qubits. Here, we demonstrate a chip-scale source of entangled optical and microwave photonic qubits. Our device platform integrates a piezo-optomechanical transducer with a superconducting resonator which is robust under optical illumination. We drive a photon-pair generation process and employ a dual-rail encoding intrinsic to our system to prepare entangled states of microwave and optical photons. We place a lower bound on the fidelity of the entangled state by measuring microwave and optical photons in two orthogonal bases. This entanglement source can directly interface telecom wavelength time-bin qubits and gigahertz frequency superconducting qubits, two well-established platforms for quantum communication and computation, respectively. Published by the American Physical Society 2024

Physics↗

FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression

Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).

SZ3↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Mutual information bounded by Fisher information

We derive a general upper bound to mutual information in terms of the Fisher information. The bound may be further used to derive a lower bound for the Bayesian quadratic cost. These two provide alternatives to other inequalities in the literature (e.g., the van Trees inequality) that are useful also for cases where the latter ones give trivial bounds. We then generalize them to the quantum case, where they bound the Holevo information in terms of the quantum Fisher information. We illustrate the usefulness of our bounds with a case study in quantum phase estimation. Here, they allow us to adapt to mutual information (useful for global strategies where the prior plays an important role), the known and highly nontrivial bounds for the Fisher information in the presence of noise. The results are also useful in the context of quantum communication, both for continuous and discrete alphabets. Published by the American Physical Society 2025

97 MATHEMATICS AND COMPUTING↗

A Bayesian Learning Approach to Wireless Outdoor Heatmap Construction using Deep Gaussian Process

We present a novel Bayesian learning approach to outdoor radio heatmap construction utilizing deep Gaussian process (GP). The proposed approach employs a two-layer hierarchy which consists of two cascaded Gaussian processes that are capable of modeling more complex input-output relations than standard single-layer Gaussian processes. Since deriving the exact model likelihood is challenging, a lower bound is optimized instead so that gradient descent-based methods can be performed to find out the optimal model parameters. Typically, inducing points are used in GPs to facilitate low-rank approximation of covariance (kernel) matrices for computation speedup. However, the inaccuracy induced by inducing points can accumulate when stacking multiple layers of GP which may hinder the performance of deep GP. Moreover, since inducing points need to be learned, having them at all layers of deep GP also incurs computational burden. To overcome the above challenges, in contrast to the canonical deep GP model, we use a modified architecture where a full standard GP resides in the first layer and inducing points are only introduced for the second layer. This modified architecture strikes a balance between model accuracy and training complexity. In the proposed model, the noise parameter of the first GP layer is also eliminated to improve the training efficiency as the noise parameter at the output of the second layer suffices to model the uncertainty in the output. The proposed approach is evaluated on real-world datasets, in the form of location-Received Signal Strength (RSS) pairs, collected from the Platform for Open Wireless Data-driven Experimental Research (POWDER) located at the campus of the University of Utah. Experiment results show that the proposed approach can achieve smaller prediction errors on various training and testing data configurations than DNN-based and GP-based methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗