Search NASA⌕ Search

SEARCH · Search NASA

Results for “quantization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Low-latency Jet Tagging for HL-LHC Using Transformer Architectures

Transformers are the state-of-the-art model architectures and widely used in application areas of machine learning. However the performance of such architectures is less well explored in the ultra-low latency domains where deployment on FPGAs or ASICs is required. Such domains include the trigger and data acquisition systems of the LHC experiments. We present a transformer-based algorithm for jet tagging built with the HGQ2 framework, which is able to produce a model with heterogeneous bitwidths for fast inference on FPGAs, as required in the trigger systems at the LHC experiments. The bitwidths are acquired during training by minimizing the total bit operations as an additional parameter. By allowing a bitwidth of zero, the model is pruned in-situ during training. Using this quantization-aware approach, our algorithm achieves state-of-the-art performance while also retaining permutation invariance which is a key property for particle physics applications. Due to the strength of transformers in representation learning, our work also serves as a stepping stone for the development of a larger foundation model for trigger applications.

Laatu, Lauri [Imperial Coll., London]↗

Evaluation of fluxon synapse device based on superconducting loops for energy efficient neuromorphic computing

With Moore’s law nearing its end due to the physical scaling limitations of CMOS technology, alternative computing approaches have gained considerable attention as ways to improve computing performance. Here, we evaluate performance prospects of a new approach based on disordered superconducting loops with Josephson-junctions for energy efficient neuromorphic computing. Synaptic weights can be stored as internal trapped fluxon states of three superconducting loops connected with multiple Josephson-junctions (JJ) and modulated by input signals applied in the form of discrete fluxons (quantized flux) in a controlled manner. The stable trapped fluxon state directs the incoming flux through different pathways with the flow statistics representing different synaptic weights. We explore implementation of matrix–vector-multiplication (MVM) operations using arrays of these fluxon synapse devices. We investigate the energy efficiency of online-learning of MNIST dataset. Our results suggest that the fluxon synapse array can provide ~100× reduction in energy consumption compared to other state-of-the-art synaptic devices. This work presents a proof-of-concept that will pave the way for development of high-speed and highly energy efficient neuromorphic computing systems based on superconducting materials.

42 ENGINEERING↗

Space-Time Finite Element Tensor Network Approach for the Time-Dependent Convection–Diffusion–Reaction Equation with Variable Coefficients

In this paper, we present a new space-time Galerkin-like method, where we treat the discretization of spatial and temporal domains simultaneously. This method utilizes a mixed formulation of the tensor-train (TT) and quantized tensor-train (QTT) (please see Section Tensor-Train Decomposition), designed for the finite element discretization (Q1-FEM) of the time-dependent convection–diffusion–reaction (CDR) equation. We reformulate the assembly process of the finite element discretized CDR to enhance its compatibility with tensor operations and introduce a low-rank tensor structure for the finite element operators. Recognizing the banded structure inherent in the finite element framework’s discrete operators, we further exploit the QTT format of the CDR to achieve greater speed and compression. Additionally, we present a comprehensive approach for integrating variable coefficients of CDR into the global discrete operators within the TT/QTT framework. The effectiveness of the proposed method, in terms of memory efficiency and computational complexity, is demonstrated through a series of numerical experiments, including a semi-linear example.

convection–diffusion–reaction equation↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

Nonlocal Effective Field Theory and Its Applications

We review recent applications of nonlocal effective field theory, particularly focusing on nonlocal chiral effective theory and nonlocal quantum electrodynamics (QED), as well as an extension of nonlocal effective theory to curved spacetime. For the chiral effective theory, we discuss the calculation of generalized parton distributions (GPDs) of the nucleon at nonzero skewness, along with the corresponding gravitational (or mechanical) form factors, within the convolution framework. In the QED application, we extend the nonlocal formulation to construct the most general nonlocal QED interaction, in which both the propagator and fundamental QED vertex are modified due to the nonlocal Lagrangian, while preserving the Ward–Green–Takahashi identities. For consistency with the modified propagator, a solid quantization is proposed, and the nonlocal QED is applied to explain the lepton g−2 anomalies without the introduction of new particles beyond the standard model. Finally, with an extension of the chiral effective action to curved spacetime, we investigate the nonlocal energy–momentum tensor and gravitational form factors of the nucleon with a nonlocal pion–nucleon interaction.

chiral effective field theory↗

Impact of Magnetic-field-driven Anisotropies on the Equation of State Probed in Neutron Star Mergers

Binary neutron star mergers can produce extreme magnetic fields, some of which can lead to strong magnetar-like remnants. While strong magnetic fields have been shown to affect the dynamics of outflows and angular momentum transport in the remnant, they can also crucially alter the properties of nuclear matter probed in the merger. In this work, we provide a first assessment of the latter, determining the strength of the pressure anisotropy caused by Landau-level quantization and the anomalous magnetic moment. To this end, we perform the first numerical relativity simulation with a magnetic polarization tensor and a magnetic-field-dependent equation of state using a new algorithm we present here, which also incorporates a mean-field dynamo model to control the magnetic field strength present in the merger remnant. Our results show that—in the most optimistic case—corrections to the anisotropy can be in excess of 10% and are potentially largest in the outer layers of the remnant. This work paves the way for a systematic investigation of these effects.

General relativity↗

Ginsparg-Wilson Hamiltonians with Improved Chiral Symmetry

We construct a family of Ginsparg-Wilson Hamiltonians with improved chiral properties, starting from a construction of Creutz-Horvath-Neuberger that provides a doubler-free Hamiltonian lattice regularization for Dirac fermions in even spacetime dimensions. We use a higher-order generalization of the Ginsparg-Wilson relation due to Fujikawa, which yields an order-$k$ Hamiltonian overlap operator for each integer $k \geq 0$, with an exactly conserved but nonquantized chiral charge that becomes quantized as $k \to \infty$. Our construction provides physical insight into how Fujikawa's higher-order Ginsparg-Wilson relation improves chiral symmetry while reproducing the anomaly, highlighting the trade-offs inherent in any Hamiltonian lattice realization of an anomalous chiral symmetry. This class of Hamiltonian lattice regularizations, with their tunable chiral symmetry properties, offers potential advantages for quantum and tensor-network simulations.

Singh, Hersh [Fermilab]↗

Fast Machine Learning for Quantum Control of Microwave Qudits on Edge Hardware

Quantum optimal control is a promising approach to improve the accuracy of quantum gates, but it relies on complex algorithms to determine the best control settings. CPU or GPU-based approaches often have delays that are too long to be applied in practice. It is paramount to have systems with extremely low delays to quickly and with high fidelity adjust quantum hardware settings, where fidelity is defined as overlap with a target quantum state. Here, we utilize machine learning (ML) models to determine control-pulse parameters for preparing Selective Number-dependent Arbitrary Phase (SNAP) gates in microwave cavity qudits, which are multi-level quantum systems that serve as elementary computation units for quantum computing. The methodology involves data generation using classical optimization techniques, ML model development, design space exploration, and quantization for hardware implementation. Our results demonstrate the efficacy of the proposed approach, with optimized models achieving low gate trace infidelity near $10^{-3}$ and efficient utilization of programmable logic resources.

Sanders, Flor [Columbia U.]↗

Rapid Inference of Logic Gate Neural Networks for Anomaly Detection in High Energy Physics

The increasing data rates and complexity of detectors at the Large Hadron Collider (LHC) necessitate fast and efficient machine learning models, particularly for rapid selection of what data to store, known as triggering. Building on recent work in differentiable logic gates, we present a public implementation of a Convolutional Differentiable Logic Gate Neural Network (CLGN). We apply this to detecting anomalies at the Level-1 Trigger at CMS using public data from the CICADA project. We demonstrate that the CLGN achieves physics performance on par with or superior to conventional quantized neural networks. We also synthesize an LGN for a Field-Programmable Gate Array (FPGA) and show highly promising FPGA characteristics, notably zero Digital Signal Processor (DSP) resource usage. This work highlights the potential of logic gate networks for high-speed, on-detector inference in High Energy Physics and beyond.

FOS: Physical sciences↗

Quantum Gravity and Laser Interferometry: Towards Observable Predictions

Understanding quantum gravity remains one of the deepest challenges in modern physics, as direct experimental access to Planck-scale effects is beyond current technological reach. However, recent theoretical advances indicate that quantum fluctuations of spacetime may produce measurable effects in precision experiments, particularly near causal horizons. This opens new avenues for testing quantum gravity phenomena through high-precision measurement techniques. This dissertation develops multiple theoretical models to characterize these effects and examines their potential observational signatures in future gravitational wave interferometers. We begin by investigating the role of quantum fluctuations in near-horizon geometries through the lens of the AdS/CFT correspondence, which provides a powerful framework for understanding the interplay between quantum field theory and general relativity via holographic principles. By modeling stochastic energy-momentum sources in Rindler-AdS spacetime, we demonstrate that vacuum fluctuations transform the Einstein equations into a Langevin-type stochastic differential equation, leading to potentially observable fluctuations in photon traversal times. Extending this approach to Minkowski spacetime, we establish a correspondence between gravitational shockwaves and fluid dynamics, showing that near-horizon perturbations satisfy an equation analogous to that governing incompressible fluids, thereby reinforcing the membrane paradigm and hydrodynamic analogies in the context of the fluid/gravity duality. Furthermore, we construct the covariant phase space of a spherically symmetric causal diamond in Minkowski spacetime, identifying two fundamental charges that govern its evolution. These results provide a foundation for quantizing causal horizons and understanding their microscopic degrees of freedom. Building upon these theoretical developments, we further examine a related stochastic phenomenon: the gravitational wave memory background arising from the cumulative memory steps produced by supermassive black hole mergers. After reviewing the standard stochastic gravitational wave background, gravitational memory effects, and BMS symmetries, we model the stochastic memory background using a Brownian motion framework. We show that while the cumulative memory background initially appears above the sensitivity curve of space-based interferometers like LISA, the realistic subtraction of individually resolvable merger events substantially suppresses the residual signal, making its detection more challenging. This highlights the critical importance of source subtraction when evaluating the detectability of gravitational memory effects. By bridging fundamental theory with experimental prospects, this dissertation contributes to the ongoing effort to uncover the quantum nature of spacetime through precision measurement techniques. Whether through detecting quantum spacetime fluctuations, gravitational memory backgrounds, or probing the symmetries of causal horizons, the pursuit of observable quantum gravity phenomena continues to expand the frontiers of both theory and experiment.

Zhang, Yiwen [Caltech] (ORCID:0000000323559416)↗

Astronomical Spectroscopy with Skipper CCDs: First Results from a Skipper CCD Focal Plane Prototype at SIFS

We present the first on-sky results from an ultra-low-readout-noise Skipper CCD focal plane prototype for the SOAR Integral Field Spectrograph (SIFS). The Skipper CCD focal plane consists of four 6k x 1k, 15 $\mu$m pixel, fully-depleted, p-channel devices that have been thinned to ~250 $\mu$m, backside processed, and treated with an anti-reflective coating. These Skipper CCDs were configured for astronomical spectroscopy, i.e., single-sample readout noise < 4.3 e- rms/pixel, the ability to achieve multi-sample readout noise $\ll$ 1 e- rms/pixel, full-well capacities ~40,000-65,000 e-, low dark current and charge transfer inefficiency (~2 x 10$^{-4}$ e-/pixel/s and 3.44 x 10$^{-7}$, respectively), and an absolute quantum efficiency of $\gtrsim$ 80% between 450 nm and 980 nm ($\gtrsim$ 90% between 600 nm and 900 nm). We optimized the readout sequence timing to achieve sub-electron noise (~0.5 e- rms/pixel) in a region of 2k x 4k pixels and photon-counting noise (~0.22 e- rms/pixel) in a region of 220 x 4k pixels, each with a readout time of $\lesssim$ 17 min. We observed two quasars (HB89 1159+123 and QSO J1621-0042) at redshift z ~ 3.5, two high-redshift galaxy clusters (CL J1001+0220 and SPT-CL J2040-4451), an emission line galaxy at z = 0.3239, a candidate member star of the Bo\"{o}tes II ultra-faint dwarf galaxy, and five CALSPEC spectrophotometric standard stars (HD074000, HD60753, HD106252, HD101452, HD200654). We present charge-quantized, photon-counting observations of the quasar HB89 1159+123 and show the detector sensitivity increase for faint spectral features. We demonstrate signal-to-noise performance improvements for SIFS observations in the low-background, readout-noise-dominated regime. We outline scientific studies that will leverage the SIFS-Skipper CCD data and new detector architectures that utilize the Skipper floating gate amplifier with faster readout times.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Hydrogen Detection Strategies to Support H2@SCALE - The NREL Sensor Laboratory

Hydrogen represents a major pathway to decarbonize and stabilize the national and international energy industry and select manufacturing markets. To facilitate the development of hydrogen markets, the US Department of Energy initiated H2@Scale to bring together stakeholders to advance affordable hydrogen production, transport, storage, and utilization to increase revenue opportunities in multiple energy sectors. One major impediment to hydrogen implementation is cost. To expedite the use of hydrogen in energy and other markets, the United States announced in 2021 the Hydrogen Shot, which seeks to reduce the cost of clean hydrogen by 80% to $1 per 1 kilogram in 1 decade ("1 1 1"). As the cost of hydrogen drops, new applications will emerge that will require unique configurations of existing equipment and infrastructure, and eventually lead to advances in the generation and utilization of hydrogen. As the hydrogen economy expands, sensors and detection methods will need to adapt to changing infrastructure demands to address the primary targets of health & safety, emissions monitoring, and process control. The NREL Sensor Laboratory is playing a pivotal role in advancing the use of hydrogen sensors and detection methodologies in each of these categories to support DOE's mission for safe and efficient utilization in emerging markets. Health & safety monitors are required to ensure that operators and facilities can react to unintended hydrogen releases, either as GH2, LH2, or as a constituent of blends (e.g., natural gas or ammonia). Current detection methodologies focus on safety applications to detect near its lower flammable limit (4 vol %), and typically include point sensors in applications such as fixed or mobile detectors (e.g., personal gas monitors). Methodologies amenable for area detection include acoustic, emerging optical imaging methods, and flame detectors. Comparable detection strategies can be utilized for emissions monitoring and quantization, however few methods can simultaneously cover both low (emissions) and high (health & safety) levels. Deployment of emission level detectors will be required to 1) reduce product loss through small but potentially significant leaks from an environmental or cost perspective, 2) reduce downtime of high demand systems by early identification of eminent system failures (leaks through pump or compressor seals indicative of impending failure), and 3) address potential emission monitoring requirements that may be set by regulating bodies. The first two points should be adopted by industry to reduce the cost-of-goods-sold. The third main category for hydrogen detection relates to process control and may be advantageous for many existing applications. Two main applications are emerging. For example, the purity requirements for hydrogen that is dispensed from refueling systems for hydrogen fuel cell electric vehicles (FCEV) is rigorously regulated by the Standard SAE J2719, which prescribes maximum allowable levels of multiple impurities in the hydrogen fuel and must be verified by a regulatory body. Hydrogen contaminant detectors (HCD) integrated to the fueling station can assure this compliance. HCDs must be able operate in 100% H2 backgrounds and be able to distinguish between multiple contaminants at low ppm to low ppb levels. Secondly, as a strategy to decarbonize the natural gas grid, there are proposals to blend hydrogen with natural gas. This blending will affect transport applications (pipeline infrastructure), stationary combustion systems (turbines), and consumer and commercial appliances. In the short-term, hydrogen levels up to 20% are proposed. Variations in the hydrogen level can have dramatic impact on the combustion process and on the potential response of safety sensors. These mixtures may be regulated so that the concentration at a delivery point must be monitored with high precision. However, routine maintenance may introduce background gases such as ambient air (with water) or maintenance gases (introduced with welding processes or adhesive outgassing.) Therefore, the detection methodology must be robust enough to recover or respond to various contaminants. Several reviews can be found in literature addressing sensing and detection technologies, including their limitations and applications. However, for most applications, limitations can be alleviated by combining various detection techniques either through system integration or implementation of machine learning methods (artificial intelligence). In this presentation, we will discuss several applications, highlight their current approach for hydrogen detection, and suggest detection strategies to supplement their limitations.

ENERGY STORAGE,HYDROGEN↗

Jamming Detection for Low-Resolution SC-FDE Systems: A Machine Learning Approach

Jammers interfere with communication between base stations (BSs) and legitimate users, leading to degradation of wireless system performance. Our study focuses on jamming detection for wideband single-carrier frequency domain equalization (SC-FDE) systems with low-resolution analog-to digital converters (ADCs). In such systems, jamming detection is challenging because traditional analytical approaches cannot be directly applied due to the delay dispersion in wideband channels and the non-linearity induced by low-resolution ADCs. We propose a machine learning (ML)-based jamming detection method that directly uses the quantized receive signals. Significantly, our ML-based detector can be integrated into existing standard frameworks, such as unique word (UW)-based SC-FDE systems, as it uses existing pilots without requiring additional pilots for jamming detection. Through numerical simulations, we show that two or more bits provide satisfactory performance compared to unquantized scenarios. Additionally, we demonstrate that using more and well-separated pilot symbols improves performance.

99 GENERAL AND MISCELLANEOUS↗

Development of a Practical Secondary Control for Hardware Microgrids: Preprint

Practical vendor-agnostic interoperability guidelines for the secondary control architecture of microgrids (MGs) with multiple grid-forming (GFM) inverter-based resources (IBRs) have not yet been developed. Therefore, this paper proposes a generic and vendor-agnostic secondary control architecture that works with all GFM IBRs and synchronous generators. This secondary control does not need to employ any additional measurement devices in the MG and uses the inherent communication systems of the GFM units, such as Modbus TCP/IP. The practical challenges of Modbus holding registers such as packet loss, quantization error, etc. and their deteriorating impacts on the secondary control action are investigated. The proposed three-stage modification of any secondary control architecture to eliminate these impacts includes 1) averaging the data read, 2) situational event-triggering of the controller, and 3) the finite iteration of the controlling action. The proposed method is validated using a laboratory hardware MG with commercial GFM units.

inverter based resources↗

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

Chitty-Venkata, Krishna Teja↗

Electron-Proton Scattering Event Generation using Structured Tokenization

Recent work such as Omnijet-$\alpha$ has demonstrated that effective tokenization combined with transformer-based architectures can produce effective foundation models for jet physics. While tokenization may help models capture generalizable event characteristics, it also introduces discretization errors that may compromise the precision required for downstream physics analyses. As the number and complexity of the particle features grow, these errors are likely to grow proportionally. In this study, we investigate new tokenization strategies to improve the application of generative transformer models to \textsc{Pythia8} simulations of electron-proton scattering at the Electron-Ion Collider. Specifically, we propose a feature-based structured tokenization approach that utilizes multiple tokens per particle, improving expressivity, while reducing the total number of unique tokens needed. We evaluate this method against grid-based binning, K-means clustering, and vector-quantized variational auto-encoders on the event simulations. Our results show that feature-based structured tokenization reduces discretization error, leading to more accurate generative modeling of particle-level events.

Goldenberg, Steven [Thomas Jefferson National Acce↗

Sub-microsecond Transformers for Jet Tagging on FPGAs

We present the first sub-microsecond transformer implementation on an FPGA achieving competitive performance for state-of-the-art high-energy physics benchmarks. Transformers have shown exceptional performance on multiple tasks in modern machine learning applications, including jet tagging at the CERN Large Hadron Collider (LHC). However, their computational complexity prohibits use in real-time applications, such as the hardware trigger system of the collider experiments up until now. In this work, we demonstrate the first application of transformers for jet tagging on FPGAs, achieving $\mathcal{O}(100)$ nanosecond latency with superior performance compared to alternative baseline models. We leverage high-granularity quantization and distributed arithmetic optimization to fit the entire transformer model on a single FPGA, achieving the required throughput and latency. Furthermore, we add multi-head attention and linear attention support to hls4ml, making our work accessible to the broader fast machine learning community. This work advances the next-generation trigger systems for the High Luminosity LHC, enabling the use of transformers for real-time applications in high-energy physics and beyond.

Laatu, Lauri [Imperial Coll., London]↗

Surrogate Neural Architecture Codesign Package (SNAC-Pack)

Neural architecture search (NAS) is a powerful approach for automating model design, but existing methods often optimize for accuracy alone or rely on proxy metrics such as bit operations (BOPs) that correlate poorly with hardware cost. This gap is particularly large for FPGA deployment, where cost is dominated by a multi-dimensional budget of lookup tables, DSPs, flip-flops, BRAM, and latency. We present the Surrogate Neural Architecture Codesign Package (SNAC-Pack), an open-source AutoML framework for hardware-aware neural architecture codesign and end-to-end FPGA deployment. SNAC-Pack runs a multi-objective global search with Optuna and NSGA-II, loading trials to a shared SQLite store that enables parallel workers across compute nodes. A hardware surrogate model outputs per-trial resource and latency estimates, avoiding the synthesis cost that would otherwise dominate the search loop. A local search stage then applies quantization-aware training (QAT) together with iterative magnitude pruning in a combined compression loop, after which the final model is synthesized to FPGA firmware via the hls4ml Python library. A YAML configuration and an optional agentic frontend let users run the pipeline on new datasets without modifying the framework. We demonstrate SNAC-Pack on jet classification at the Large Hadron Collider and superconducting qubit readout, discovering compact architectures that match or exceed strong baselines on the task metric while reducing FPGA resource utilization and, in the qubit readout case, reducing the design space exploration process from months of manual fine-tuning to hours of automated search.

Weitz, Jason [UC, San Diego]↗