Search NASA⌕ Search

SEARCH · Search NASA

Results for “quantization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Regime Characterization of Offshore Wind Resource Using Unsupervised Learning

Predictability of wind resource conditions is critical for offshore wind design and operations. While many studies of extreme wind conditions focus on specific events such as low-level jets or ramps, these rely on threshold definitions that limit generality. Here we present a data-driven framework that combines principal component analysis (PCA), self-organizing maps (SOM), and k-means clustering to classify wind resource conditions as typical and anomalous from climatological data. Anomalies are defined not by fixed thresholds but by flagging samples located far from SOM node centers inside the baseline SOM structure. This reframes extremes as rare ebents and hence, likely difficult to anticipate by numerical weather prediction models. We applied this approach to 23 years (2000–2022) of hourly profiles from the NOW-23 hindcast model at the Humboldt Wind Energy Area. Classification is conducted on a feature space consisting of 10 m wind speed and direction, bulk shear and veer across 30–270 m, and a low-level jet index. Dimensionality reduction is achieved through PC. A 2 × 3 OM lattice trained on the PCA vectors identified six baseline regimes spanning weak to strong flow states. High quantization-error profiles are identified and re-clustered into four anomalous regimes. The baseline regimes exhibited clear seasonal and diurnal cycles. Meanwhile, the anomalous regimes represented <10 % of all hours but showed distinct combinations of speed, shear, and veer, when compared to the baseline regimes. Anomalous regimes are typically short-lived (~few hours), yet their transitions can lead to hub-height wind changes of −18 to +9 m s -1 . For a representative 15 MW turbine, these shifts imply rapid swings in capacity factor from near-full output to negligible generation. Validation with lidar buoy data showed 51% agreement in SOM labels across ~6,000 overlapping hours, with most mismatches confined to adjacent speed classes. HRRR comparisons further revealed that anomalous regimes were disproportionately associated with forecast biases exceeding 5 m s -1 . Together, these results reframe extremes in offshore wind from absolute maxima or minima to weather states that are difficult to anticipate from models.

17 WIND ENERGY↗

Scalable workflow for evaluating and optimizing large language models

This work describes the improved workflow for evaluating open-source large language models (LLMs) for trustworthiness. The workflow facilitates the acquisition of LLMs, the generation of LLM responses, and the evaluation of the responses for their trustworthiness. As a use case, the workflow is employed to evaluate dense, quantized, and pruned Meta Llama3.1 LLMs for their truthfulness. The outcome of the project could set the stage for understanding and developing trustworthy models in the future projects.

97 MATHEMATICS AND COMPUTING↗

Energy-efficient, Large-scale Molecular Dynamics Simulations via Hardware- and Algorithm-level Optimization

This work aims to develop a framework for energy-efficient computing that will enable molecular dynamics (MD) simulations of large-scale phenomena with atomic precision and simultaneously remove computational bottlenecks limiting the speed of MD simulations. We seek to implement such an approach through the development of surrogate models for the interatomic force calculation combined with the use of mixed numerical precision formats. For a model system of neutral atoms (only pairwise interactions), significant force calculation efficiency improvements were achieved, without detrimental effects on atomic structures or average energies, using single precision, by developing a surrogate model (deep neural network), and by quantizing this surrogate model. For a model system of charged atoms, the reciprocal-space calculation of electrostatic interactions was identified as the main bottleneck, and the development of a surrogate model should be pursued to achieve an estimated one-order-of-magnitude additional speedup.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Electron-Proton Scattering Event Generation using Structured Tokenization

Recent work such as Omnijet-$\alpha$ has demonstrated that effective tokenization combined with transformer-based architectures can produce effective foundation models for jet physics. While tokenization may help models capture generalizable event characteristics, it also introduces discretization errors that may compromise the precision required for downstream physics analyses. As the number and complexity of the particle features grow, these errors are likely to grow proportionally. In this study, we investigate new tokenization strategies to improve the application of generative transformer models to \textsc{Pythia8} simulations of electron-proton scattering at the Electron-Ion Collider. Specifically, we propose a feature-based structured tokenization approach that utilizes multiple tokens per particle, improving expressivity, while reducing the total number of unique tokens needed. We evaluate this method against grid-based binning, K-means clustering, and vector-quantized variational auto-encoders on the event simulations. Our results show that feature-based structured tokenization reduces discretization error, leading to more accurate generative modeling of particle-level events.

Goldenberg, Steven [Thomas Jefferson National Acce↗

UNLOQ: UNderstanding coherence in Light-matter interfaces for Quantum Science (Final Technical Report)

The general goal of this project is to prepare next-generation quantum systems for novel quantum information science applications. Decades of research on quantum optics have provided revolutionary systems for manipulating atomic and photonic quantum coherence in transformative ways. Our hypothesis is that the next-generation quantum systems will come from nano-molecular quantum optics. Specifically, we are interested in coupling the electronic states of molecules or nanoparticles to the quantized radiation field inside an optical cavity to create a set of new photon-matter hybrid excitations, called polaritons. As opposed to atoms, the vibrational modes of molecules and nanoparticles provide new degrees of freedom to mediate the quantum transduction between electronic and photonic states, offering new ways to tune and ultimately control the quantum coherence of the integrated system. In this project we are interested in designing, fabricating, and characterizing the polaritons that arise from coupling CdSe nanoplatelets (NPLs) to a Fabry-Pérot optical cavity. Specific attention was paid to parameters of the system (cavity mode volume, quality factor, etc.) that would maximize the collective coupling strength. Through a combined theoretical and experimental approach, we were able to make important contributions to understanding NPL exciton-polariton photophysics including how to use cavity loss as a tunable parameter in order to manipulate the populations of the upper and lower polariton states. Finally, using state-of-the-art theoretical methods, we were able to show how the collective coupling of many molecular excitons in a cavity can protect polariton coherence from vibrationally-induced decoherence. In particular, for NPL-cavity polaritonic systems, the quantum coherence can be extended by over an order of magnitude at room temperature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Low-latency Jet Tagging for HL-LHC Using Transformer Architectures

Transformers are the state-of-the-art model architectures and widely used in application areas of machine learning. However the performance of such architectures is less well explored in the ultra-low latency domains where deployment on FPGAs or ASICs is required. Such domains include the trigger and data acquisition systems of the LHC experiments. We present a transformer-based algorithm for jet tagging built with the HGQ2 framework, which is able to produce a model with heterogeneous bitwidths for fast inference on FPGAs, as required in the trigger systems at the LHC experiments. The bitwidths are acquired during training by minimizing the total bit operations as an additional parameter. By allowing a bitwidth of zero, the model is pruned in-situ during training. Using this quantization-aware approach, our algorithm achieves state-of-the-art performance while also retaining permutation invariance which is a key property for particle physics applications. Due to the strength of transformers in representation learning, our work also serves as a stepping stone for the development of a larger foundation model for trigger applications.

Laatu, Lauri [Imperial Coll., London]↗

Evaluation of fluxon synapse device based on superconducting loops for energy efficient neuromorphic computing

With Moore’s law nearing its end due to the physical scaling limitations of CMOS technology, alternative computing approaches have gained considerable attention as ways to improve computing performance. Here, we evaluate performance prospects of a new approach based on disordered superconducting loops with Josephson-junctions for energy efficient neuromorphic computing. Synaptic weights can be stored as internal trapped fluxon states of three superconducting loops connected with multiple Josephson-junctions (JJ) and modulated by input signals applied in the form of discrete fluxons (quantized flux) in a controlled manner. The stable trapped fluxon state directs the incoming flux through different pathways with the flow statistics representing different synaptic weights. We explore implementation of matrix–vector-multiplication (MVM) operations using arrays of these fluxon synapse devices. We investigate the energy efficiency of online-learning of MNIST dataset. Our results suggest that the fluxon synapse array can provide ~100× reduction in energy consumption compared to other state-of-the-art synaptic devices. This work presents a proof-of-concept that will pave the way for development of high-speed and highly energy efficient neuromorphic computing systems based on superconducting materials.

42 ENGINEERING↗

Space-Time Finite Element Tensor Network Approach for the Time-Dependent Convection–Diffusion–Reaction Equation with Variable Coefficients

In this paper, we present a new space-time Galerkin-like method, where we treat the discretization of spatial and temporal domains simultaneously. This method utilizes a mixed formulation of the tensor-train (TT) and quantized tensor-train (QTT) (please see Section Tensor-Train Decomposition), designed for the finite element discretization (Q1-FEM) of the time-dependent convection–diffusion–reaction (CDR) equation. We reformulate the assembly process of the finite element discretized CDR to enhance its compatibility with tensor operations and introduce a low-rank tensor structure for the finite element operators. Recognizing the banded structure inherent in the finite element framework’s discrete operators, we further exploit the QTT format of the CDR to achieve greater speed and compression. Additionally, we present a comprehensive approach for integrating variable coefficients of CDR into the global discrete operators within the TT/QTT framework. The effectiveness of the proposed method, in terms of memory efficiency and computational complexity, is demonstrated through a series of numerical experiments, including a semi-linear example.

convection–diffusion–reaction equation↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Nonlocal Effective Field Theory and Its Applications

We review recent applications of nonlocal effective field theory, particularly focusing on nonlocal chiral effective theory and nonlocal quantum electrodynamics (QED), as well as an extension of nonlocal effective theory to curved spacetime. For the chiral effective theory, we discuss the calculation of generalized parton distributions (GPDs) of the nucleon at nonzero skewness, along with the corresponding gravitational (or mechanical) form factors, within the convolution framework. In the QED application, we extend the nonlocal formulation to construct the most general nonlocal QED interaction, in which both the propagator and fundamental QED vertex are modified due to the nonlocal Lagrangian, while preserving the Ward–Green–Takahashi identities. For consistency with the modified propagator, a solid quantization is proposed, and the nonlocal QED is applied to explain the lepton g−2 anomalies without the introduction of new particles beyond the standard model. Finally, with an extension of the chiral effective action to curved spacetime, we investigate the nonlocal energy–momentum tensor and gravitational form factors of the nucleon with a nonlocal pion–nucleon interaction.

chiral effective field theory↗

Impact of Magnetic-field-driven Anisotropies on the Equation of State Probed in Neutron Star Mergers

Binary neutron star mergers can produce extreme magnetic fields, some of which can lead to strong magnetar-like remnants. While strong magnetic fields have been shown to affect the dynamics of outflows and angular momentum transport in the remnant, they can also crucially alter the properties of nuclear matter probed in the merger. In this work, we provide a first assessment of the latter, determining the strength of the pressure anisotropy caused by Landau-level quantization and the anomalous magnetic moment. To this end, we perform the first numerical relativity simulation with a magnetic polarization tensor and a magnetic-field-dependent equation of state using a new algorithm we present here, which also incorporates a mean-field dynamo model to control the magnetic field strength present in the merger remnant. Our results show that—in the most optimistic case—corrections to the anisotropy can be in excess of 10% and are potentially largest in the outer layers of the remnant. This work paves the way for a systematic investigation of these effects.

General relativity↗

Ginsparg-Wilson Hamiltonians with Improved Chiral Symmetry

We construct a family of Ginsparg-Wilson Hamiltonians with improved chiral properties, starting from a construction of Creutz-Horvath-Neuberger that provides a doubler-free Hamiltonian lattice regularization for Dirac fermions in even spacetime dimensions. We use a higher-order generalization of the Ginsparg-Wilson relation due to Fujikawa, which yields an order-$k$ Hamiltonian overlap operator for each integer $k \geq 0$, with an exactly conserved but nonquantized chiral charge that becomes quantized as $k \to \infty$. Our construction provides physical insight into how Fujikawa's higher-order Ginsparg-Wilson relation improves chiral symmetry while reproducing the anomaly, highlighting the trade-offs inherent in any Hamiltonian lattice realization of an anomalous chiral symmetry. This class of Hamiltonian lattice regularizations, with their tunable chiral symmetry properties, offers potential advantages for quantum and tensor-network simulations.

Singh, Hersh [Fermilab]↗

Fast Machine Learning for Quantum Control of Microwave Qudits on Edge Hardware

Quantum optimal control is a promising approach to improve the accuracy of quantum gates, but it relies on complex algorithms to determine the best control settings. CPU or GPU-based approaches often have delays that are too long to be applied in practice. It is paramount to have systems with extremely low delays to quickly and with high fidelity adjust quantum hardware settings, where fidelity is defined as overlap with a target quantum state. Here, we utilize machine learning (ML) models to determine control-pulse parameters for preparing Selective Number-dependent Arbitrary Phase (SNAP) gates in microwave cavity qudits, which are multi-level quantum systems that serve as elementary computation units for quantum computing. The methodology involves data generation using classical optimization techniques, ML model development, design space exploration, and quantization for hardware implementation. Our results demonstrate the efficacy of the proposed approach, with optimized models achieving low gate trace infidelity near $10^{-3}$ and efficient utilization of programmable logic resources.

Sanders, Flor [Columbia U.]↗

Rapid Inference of Logic Gate Neural Networks for Anomaly Detection in High Energy Physics

The increasing data rates and complexity of detectors at the Large Hadron Collider (LHC) necessitate fast and efficient machine learning models, particularly for rapid selection of what data to store, known as triggering. Building on recent work in differentiable logic gates, we present a public implementation of a Convolutional Differentiable Logic Gate Neural Network (CLGN). We apply this to detecting anomalies at the Level-1 Trigger at CMS using public data from the CICADA project. We demonstrate that the CLGN achieves physics performance on par with or superior to conventional quantized neural networks. We also synthesize an LGN for a Field-Programmable Gate Array (FPGA) and show highly promising FPGA characteristics, notably zero Digital Signal Processor (DSP) resource usage. This work highlights the potential of logic gate networks for high-speed, on-detector inference in High Energy Physics and beyond.

FOS: Physical sciences↗

Quantum Gravity and Laser Interferometry: Towards Observable Predictions

Understanding quantum gravity remains one of the deepest challenges in modern physics, as direct experimental access to Planck-scale effects is beyond current technological reach. However, recent theoretical advances indicate that quantum fluctuations of spacetime may produce measurable effects in precision experiments, particularly near causal horizons. This opens new avenues for testing quantum gravity phenomena through high-precision measurement techniques. This dissertation develops multiple theoretical models to characterize these effects and examines their potential observational signatures in future gravitational wave interferometers. We begin by investigating the role of quantum fluctuations in near-horizon geometries through the lens of the AdS/CFT correspondence, which provides a powerful framework for understanding the interplay between quantum field theory and general relativity via holographic principles. By modeling stochastic energy-momentum sources in Rindler-AdS spacetime, we demonstrate that vacuum fluctuations transform the Einstein equations into a Langevin-type stochastic differential equation, leading to potentially observable fluctuations in photon traversal times. Extending this approach to Minkowski spacetime, we establish a correspondence between gravitational shockwaves and fluid dynamics, showing that near-horizon perturbations satisfy an equation analogous to that governing incompressible fluids, thereby reinforcing the membrane paradigm and hydrodynamic analogies in the context of the fluid/gravity duality. Furthermore, we construct the covariant phase space of a spherically symmetric causal diamond in Minkowski spacetime, identifying two fundamental charges that govern its evolution. These results provide a foundation for quantizing causal horizons and understanding their microscopic degrees of freedom. Building upon these theoretical developments, we further examine a related stochastic phenomenon: the gravitational wave memory background arising from the cumulative memory steps produced by supermassive black hole mergers. After reviewing the standard stochastic gravitational wave background, gravitational memory effects, and BMS symmetries, we model the stochastic memory background using a Brownian motion framework. We show that while the cumulative memory background initially appears above the sensitivity curve of space-based interferometers like LISA, the realistic subtraction of individually resolvable merger events substantially suppresses the residual signal, making its detection more challenging. This highlights the critical importance of source subtraction when evaluating the detectability of gravitational memory effects. By bridging fundamental theory with experimental prospects, this dissertation contributes to the ongoing effort to uncover the quantum nature of spacetime through precision measurement techniques. Whether through detecting quantum spacetime fluctuations, gravitational memory backgrounds, or probing the symmetries of causal horizons, the pursuit of observable quantum gravity phenomena continues to expand the frontiers of both theory and experiment.

Zhang, Yiwen [Caltech] (ORCID:0000000323559416)↗