Search NASASearch

SEARCH · Search NASA

Results for “quantization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Bridging the Gap Between LLMs and LNS with Dynamic Data Format and Architecture Codesign

Deep neural networks (DNNs) have achieved tremendous success in the past few years. However, their training and inference demand exceptional computational and memory resources. Quantization has been shown as an effective approach to mitigate the cost, with the mainstream data types reduced from FP32 to FP16/BF16 and recently FP8 in the latest NVIDIA H100 GPUs. With increasingly aggressive quantization, however, the conventional floating-point formats suffer from limited precision in representing numbers around zero. Recently, NVIDIA demonstrated the potential of using a Logarithmic Number System (LNS) for the next generation of tensor cores. While LNS mitigates the hurdles in representing small numbers, in this work we observed a mismatch between LNS and the emerging Large Language Models (LLM), where LLM exhibits significant outliers when directly adopting the LNS format. In this paper, we present a data-format/architecture codesign to bright this gap. On the format side, we propose a dynamic LNS format to flexibly represent outliers at a higher precision, by exploiting asymmetry in the LNS representation and identifying outliers through a per-vector basis. On the architecture side, for demonstration, we realize the dynamic LNS format in a systolic array, which can handle the irregularity of the outliers at runtime. We implement our approach on an Alveo U280 FPGA as a prototype. Experimental results show that our design can effectively handle the outliers and resolve the mismatch between LNS and LLM, contributing to an accuracy improvement of 15.4% and 16% over the floating-point and the original LNS baselines, using four state-of-the-art LLM models. Our observation and design lay a solid foundation for the large-scale adoption of the LNS format in the next-generation deep learning hardware.

Haghi, Pouya

Interpreting and Accelerating Transformers for Jet Tagging

Attention-based transformers are ubiquitous in machine learning applications from natural language processing to computer vision. In high energy physics, one central application is to classify collimated particle showers in colliders based on the particle of origin, known as jet tagging. In this work, we study the interpretatbility and prospects for acceleration of Particle Transformer (ParT), a state-of-the-art model, leverages particle-level attention to improve jet-tagging performance. We analyzing ParT's attention maps and particle-pair correlations in the eta-phi plane, revealing intriguing features, such as a binary attention pattern that identifies critical substructure in jets. These insights enhance our understanding of the model's internal workings and learning process and hint at ways to improve its efficiency. Along these lines, we also explore low-rank attention, attention alternatives, and dynamic quantization to accelerate transformers for jet tagging. With quantization, we achieve a 50% reduction in model size and a 10% increase in inference speed without compromising accuracy. These combined efforts enhance both the performance and the interpretability of transformers in high-energy physics, opening avenues for more efficient and physics-driven model designs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Quantum Information for Fusion Energy Sciences (Final Technical Report)

The simulation of plasma dynamics is a critical area of Fusion Energy Sciences (FES) due to it’s usefulness in predicting, controlling, and confining plasmas in the context of potential fusion reactors. The simulation of plasmas is a computationally difficult problem in both classical and quantum physics, motivating investigation into the potential of quantum computers to simulate these systems. This project took several concrete steps towards this goal by developing tools for improving the control, characterization, and calibration of quantum gates on a superconducting quantum computer, developing error suppression and mitigation tools to reduce errors on the quantum computer, and utilizing these advancements to simulate reduced models of plasma dynamics on the quantum computer. In order to efficiently simulate plasma physics, an optimal control method which synthesizes, directly at the pulse level, any quantum gate on qubit and qutrit systems was developed. Using four superconducting transmon quantum processors at Rigetti and LLNL, it was demonstrated that any arbitrary quantum gate on qubits and qutrits could be implemented with high fidelity, leading to a significantly reduced length of a gate sequence. A problem of interest in FES is the nonlinear optical process of laser pulse compression within a plasma. Since quantum physics is linear, simulating nonlinear operations is not naturally feasible on a quantum computer, however it is possible to simulated a quantized version of the nonlinear process. A quantization approach to convert nonlinear wave-wave interaction problems to Hamiltonian simulation problems was developed and demonstrated using two qubits on a Rigetti device. In this experiment, a number of error suppression and mitigation techniques were investigated to determine how best to utilize the finite quantum resources. This study provides an example of how plasma problems may be solved on near-term, noisy quantum computing platforms and identified a promising set of techniques. Building on the insights of these experiments, the investigation turned to linear electron-plasma wave physics. A connection was identified between a local one-dimensional lattice spin model and linear wave phenomena, allowing a plasma physics problem to be efficiently mapped to the quantum computer. In this framework, reflection and transmission of plasma waves at a sharp boundary was studied, as well as the propagation of waves through an inhomogeneous plasma medium. In addition to the suite of error suppression and mitigation techniques developed, this experiment introduced the use of a digital-analog gate scheme designed to efficiently simulate the plasma Hamiltonian. With hardware available at the conclusion of the project, simulation at the scale of 9 qubits and 15 timesteps (60 entangling layers) was achieved.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Quantum surrogate models for uncertainty quantification

Surrogate models are a critical ingredient to computation-based design and validation of many DOE mission-relevant physical systems. When first-principles computation of properties of a physical systems becomes pro hibitive, surrogate models are the only path towards achieving tasks such as uncertainty quantification (UQ), exploration of design space, and validation of design choices. In this project we have developed and demonstrated a new surro gate modeling paradigm for complex models that is data-driven, non-intrusive, and has the potential to be versatile and equipped with performance guaran tees. This combination of features is absent in existing surrogate modeling tools. The framework we have developed in this project exploits a quantum-classical correspondence to establish a quantum system that mimics the dynamics of the classical Hamiltonian system from which data in the form of temporal snapshots is provided. Since quantum dynamics propagates distributions over observables, the framework is naturally suited to propagation of epistemic uncertainties in the form of distributions over initial state and parametric uncertainties. In this project, we take the first step in establishing this novel framework by deriving a quantization and de-quantization procedure, demonstrating the accuracy of the quantum surrogate models these define using two model systems, and defining the next steps in maturing the framework towards a tool applicable to Sandia mission-relevant problems.

97 MATHEMATICS AND COMPUTING

Boundedness regions of discrete-time dynamic systems

Techniques for obtaining quantitative information about boundedness properties are developed and applied to the sampled-data control of satellite attitude with quantization. Relevant stability concepts are introduced as a series of definitions, and interrelationships between various definitions are discussed. The boundedness regions are estimated by means of quadratic Liapunov functions, and a sufficient condition for the existence of a boundedness region is given for a certain class of systems. A quadratic Liapunov function is applied to the Lur'e-Postinkov class of systems, where the linear part of the system is not asymptotically stable and the quantizer represents the nonlinear characteristic. A numerical calculation of the region of boundedness estimates is performed for satellite attitude control and is compared with simulation results. It is tentatively concluded that the Liapunov results may be good and that simulation results may be difficult to interpret and time-consuming to generate. The Lur'e-based technique yields estimates of regions of absolute boundedness, but at the cost of greater analytical complexity.

Siljak, D.

Viterbi decoding for satellite and space communication.

Convolutional coding and Viterbi decoding, along with binary phase-shift keyed modulation, is presented as an efficient system for reliable communication on power limited satellite and space channels. Performance results, obtained theoretically and through computer simulation, are given for optimum short constraint length codes for a range of code constraint lengths and code rates. System efficiency is compared for hard receiver quantization and 4 and 8 level soft quantization. The effects on performance of varying of certain parameters relevant to decoder complexity and cost are examined. Quantitative performance degradation due to imperfect carrier phase coherence is evaluated and compared to that of an uncoded system. As an example of decoder performance versus complexity, a recently implemented 2-Mbit/sec constraint length 7 Viterbi decoder is discussed. Finally a comparison is made between Viterbi and sequential decoding in terms of suitability to various system requirements.

Heller, J. A.

Permutation codes for sources.

Source encoding techniques based on permutation codes are investigated. For a broad class of distortion measures it is shown that optimum encoding of a source permutation code is easy to instrument even for very long block lengths. Also, the nonparametric nature of permutation encoding is well suited to situations involving unknown source statistics. For the squared-error distortion measure a procedure for generating good permutation codes of a given rate and block length is described. The performance of such codes for a memoryless Gaussian source is compared both with the rate-distortion function bound and with the performance of various quantization schemes. The comparison reveals that permutation codes are asymptotically ideal for small rates and perform as well as the best entropy-coded quantizers presently known for intermediate rates. They can be made to compare favorably at high rates, too, provided the coding delay associated with extremely long block lengths is tolerable.

Berger, T.

The use of minimum order state observers in digital flight-control systems.

This paper deals with the problem of selecting the 'arbitrary' design parameters of digital state observers when they are being used as a part of a digital flight-control system. A cost index is developed which indicates the output noise caused by input quantization due to analog-to-digital conversion. The cost index assumes that the input quantization error is uniformly distributed over the least-significant-bit of the conversion. Formulas relating the cost index to the observer design parameters are presented. The cost index is minimized with respect to the design parameters using a conjugate gradient algorithm. An example of the theory is presented in which a digital observer is designed so that a satisfactory digital flight-control system is obtained starting from an unacceptable one.

Montgomery, R. C.

Strapdown system performance optimization test evaluations (SPOT), volume 1

A three axis inertial system was packaged in an Apollo gimbal fixture for fine grain evaluation of strapdown system performance in dynamic environments. These evaluations have provided information to assess the effectiveness of real-time compensation techniques and to study system performance tradeoffs to factors such as quantization and iteration rate. The strapdown performance and tradeoff studies conducted include: (1) Compensation models and techniques for the inertial instrument first-order error terms were developed and compensation effectivity was demonstrated in four basic environments; single and multi-axis slew, and single and multi-axis oscillatory. (2) The theoretical coning bandwidth for the first-order quaternion algorithm expansion was verified. (3) Gyro loop quantization was identified to affect proportionally the system attitude uncertainty. (4) Land navigation evaluations identified the requirement for accurate initialization alignment in order to pursue fine grain navigation evaluations.

Blaha, R. J.

Search properties of some sequential decoding algorithms.

Sequential decoding procedures are studied in the context of selecting a path through a tree. Several algorithms are considered, and their properties are compared. It is shown that the stack algorithm introduced by Zigangirov (1966) and by Jelinek (1969) is essentially equivalent to the Fano algorithm with regard to the set of nodes examined and the path selected, although the description, implementation, and action of the two algorithms are quite different. A modified Fano algorithm is introduced, in which the quantizing parameter is eliminated. It can be inferred from limited simulation results that, at least in some applications, the new algorithm is computationally inferior to the old. However, it is of some theoretical interest since the conventional Fano algorithm may be considered to be a quantized version of it.

Geist, J. M.

Information preserving coding for multispectral data

A general formulation of the data compression system is presented. A method of instantaneous expansion of quantization levels by reserving two codewords in the codebook to perform a folding over in quantization is implemented for error free coding of data with incomplete knowledge of the probability density function. Results for simple DPCM with folding and an adaptive transform coding technique followed by a DPCM technique are compared using ERTS-1 data.

Duan, J. R.

Techniques for decoding speech phonemes and sounds: A concept

Techniques studied involve conversion of speech sounds into machine-compatible pulse trains. (1) Voltage-level quantizer produces number of output pulses proportional to amplitude characteristics of vowel-type phoneme waveforms. (2) Pulses produced by quantizer of first speech formants are compared with pulses produced by second formants.

Lokerson, D. C.

Research study on stabilization and control. Modern sampled-data control theory. Analysis and design of the digital large space telescope system

Dynamic modeling of a low cost single axis LST system and the pointing stability of that system are investigated. The effects of nonlinear friction of the bearings of the reaction wheels, quantization, and sensor noise on the pointing error are covered. In addition, data are given on self sustained oscillations of the system induced by quantization and methods of evaluating attitude error of the digital LST.

Kuo, B. C.

Research study on stabilization and control. Modern sampled-date control theory. Stability analysis of the low cost large space telescope system

A mathematical model employing Fourier series is used to show quantization and reaction wheel friction nonlinearity in a telescope system for use in space. Block diagrams are used to illustrate the system. A discrete describing function of a quantizer also is given, and input and output signal waveforms, with illustrative examples, are shown.

Kuo, B. C.

Formulation of the information capacity of the optical-mechanical line-scan imaging process

An expression for the information capacity of the optical-mechanical line-scan imaging process is derived which includes the effects of blurring of spatial, photosensor noise, aliasing, and quantization. Both the information capacity for a fixed data density and the information efficiency (the ratio of information capacity to data density) exhibit a distinct single maximum when displayed as a function of sampling rate, and the location of this maximum was determined by the system frequency-response shape, signal-to-noise ratio, and quantization interval.

Huck, F. O.

Performance of noncoherent MFSK channels with coding

Computer simulation of data transmission over a noncoherent channel with predetection signal-to-noise ratio of 1 shows that convolutional coding can reduce the energy requirement by 4.5 dB at a bit error rate of 0.001. The effects of receiver quantization and choice of number of tones are analyzed; nearly optimum performance is attained with eight quantization levels and sixteen tones at predetection S/N ratio of 1. The effects of changing predetection S/N ratio are also analyzed; for lower predetection S/N ratio, accurate extrapolations can be made from the data, but for higher values, the results are more complicated. These analyses will be useful in designing telemetry systems when coherence is limited by turbulence in the signal propagation medium or oscillator instability.

Butman, S. A.

Optical-mechanical line-scan imaging process - Its information capacity and efficiency

Optical-mechanical line-scan techniques have been applied to earth satellite multispectral imaging systems. The capability of the imaging system is generally assessed by its information capacity. An approach based on information theory is developed to formulate the capacity of the line-scan process. Included are the effects of blurring of spatial detail, photosensor noise, aliasing, and quantization. The information efficiency is shown to be dependent on sampling rate, system frequency response shape, SNR, and quantization interval.

Huck, F. O.

Quantum principles and free particles

The quantum principles that establish the energy levels and degeneracies needed to evaluate the partition functions are explored. The uncertainty principle is associated with the dual wave-particle nature of the model used to describe quantized gas particles. The Schroedinger wave equation is presented as a generalization of Maxwell's wave equation; the former applies to all particles while the Maxwell equation applies to the special case of photon particles. The size of the quantum cell in phase space and the representation of momentum as a space derivative operator follow from the uncertainty principle. A consequence of this is that steady-state problems that are space-time dependent for the classical model become only space dependent for the quantum model and are often easier to solve. The partition function is derived for quantized free particles and, at normal conditions, the result is the same as that given by the classical phase integral. The quantum corrections that occur at very low temperatures or high densities are derived. These corrections for the Einstein-Bose gas qualitatively describe the condensation effects that occur in liquid helium, but are unimportant for most practical purposes otherwise. However, the corrections for the Fermi-Dirac gas are important because they quantitatively describe the behavior of high-density conduction electron gases in metals and explain the zero point energy and low specific heat exhibited in this case.

Source record