Search NASA⌕ Search

SEARCH · Search NASA

Results for “quantization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

prune_quant_vit

This tool provides codes to prune and quantize vision transformers (ViTs). The tool will allow developers and researchers to speed up inference of ViTs and deploy them on CPUs.

Bhardwaj, Kshitij↗

Sentinel

Network intrusion detection systems (NIDS) are commonplace in network security but they frequently employ algorithms that are computational demanding requiring hardware and software with significant power requirements. Two examples of such resource-intensive algorithms used for network security are regular expression matching and broader signature pattern matching which are commonly used in deep packet inspection (DPI). Network security algorithms that have large power requirements may be a challenge for low-power internet-of-things (IoT) environments, which generally lack the power resources to implement complex security measures like computationally expensive DPI at the edge. Furthermore, IoT environments incorporating 5G standalone networks have network latency constraints beyond just power that make DPI at the edge even more difficult. Programmable logic is ideally suited for machine learning inference for DPI because of its deep instruction level parallelism and single-cycle memory access. Machine learning approaches for DPI have been explored before using the programmable logic of field programmable gate arrays (FPGA) as a potential solution for NIDS approaches that would be power-suitable for IoT. However, those previous programmable logic NIDS approaches utilize either a supervised or unsupervised learning model. Sentinel utilizes the ensemble of these two machine learning approaches known as a semi-supervised approach which has shown promise in NIDS implementations. Sentinel provides a programmable logic implementation of a semi-supervised approach for DPI which operates at much lower power and latency than a GPU implementation with negligible loss of accuracy due to quantization through a logistic regressor.

Anderson, MatthewW [Idaho National Laboratory (INL↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

CODARcode/MGARD

MGARD is a software providing error-controlled lossy compression and data refactoring based on multi-grid theories. It transforms floating-point scientific data into a multilevel representation, followed by quantization and lossless encoding processes, resulting in a self-describing compressed buffer. It supports diverse data topologies, error control norms, and computing architectures.

Chen, Jieyang [University of Oregon]↗

PCA -BASED COMPRESSOR HARDWARE DESIGN GENERATOR IN CHISEL

SF-25-077 This repository includes a hardware generator written in Chisel that generates lossy compression hardware designs based on principal component analysis (PCA), along with a flexible testbench. It also includes a Python tool for evaluating the accuracy loss of integer quantization.

Kazutomo, Yoshi [Argonne National Laboratory (ANL)↗

Interference dislocations adjacent to emission spot

We studied interference dislocations (forks) adjacent to an emission spot in an interference pattern. We observed the adjacent interference dislocations in emission of excitons in a monolayer transition metal dichalcogenide and in emission of spatially indirect excitons, also known as interlayer excitons, in a van der Waals heterostructure. Our simulations show that the adjacent interference dislocations appear due to the moiré effect in combined interference patterns produced by constituent parts of the emission spot. In contrast to interference dislocations in coherent states, such as interference dislocations due to quantized vortices in condensates, the appearance of adjacent interference dislocations does not require coherence between the parts of the emission spot, indicating that interference dislocations can be observed in a classical system. Finally, we show that the interference dislocations in classical systems can appear in interference images for various spatially modulated emission patterns.

Leonard, J. R. [Univ. of California, San Diego, CA↗

Asymptotic structure of higher dimensional Yang-Mills theory

Using the covariant phase space formalism, we construct the phase space for non-Abelian gauge theories in (d+2)-dimensional Minkowski spacetime for any d ≥ 2, including the edge modes that symplectically pair to the low energy degrees of freedom of the gauge field. Despite the fact that the symplectic form in odd and even-dimensional spacetimes appear ostensibly different, we demonstrate that both cases can be treated in a unified manner by utilizing the shadow transform. Upon quantization, we recover the algebra of the vacuum sector of the Hilbert space and derive a Ward identity that implies the leading soft gluon theorem in (d+2)-dimensional spacetime.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A microscopic realization of dS$_3$

We propose a precise duality between pure de Sitter quantum gravity in 2+1 2 + 1 dimensions and a double-scaled matrix integral. This duality unfolds in two distinct aspects. First, by carefully quantizing the gravitational phase space, we arrive at a novel proposal for the quantum state of the universe at future infinity. We compute cosmological correlators of massive particles in the universe specified by this wavefunction. Integrating these correlators over the metric at future infinity yields gauge-invariant observables, which are identified with the string amplitudes of the complex Liouville string [S. Collier et al., arXiv: 2409.17246]. This establishes a direct connection between integrated cosmological correlators and the resolvents of the matrix integral dual to the complex Liouville string, thereby demonstrating one aspect of the dS _3 3 /matrix integral duality. The second aspect concerns the cosmological horizon of the dS static patch and the Gibbons-Hawking entropy it is conjectured to encode. We show that this entropy can be reproduced exactly by counting the entries of the matrix.

Collier, Scott (ORCID:0000000286476653)↗

Zero-bias conductance peaks at zero applied magnetic field due to stray fields from integrated micromagnets in hybrid nanowire quantum dots

Many recipes for realizing topological superconductivity rely on broken time-reversal symmetry, which is often attained by applying a substantial external magnetic field. Alternatively, using magnetic materials can offer advantages through low-field operation and design flexibility on the nanoscale. Mechanisms for lifting spin degeneracy include exchange coupling, spin-dependent scattering, spin injection – all requiring direct contact between the bulk or induced superconductor and a magnetic material. Here, we implement locally broken time-reversal symmetry through dipolar coupling from nearby micromagnets to superconductor-semiconductor hybrid nanowire devices. Josephson supercurrent is hysteretic due to micromagnets switching. At or around zero external magnetic field, we observe an extended presence of Andreev bound states near zero voltage bias. We also show a zero-bias peak plateau of a non-quantized value. Our findings largely reproduce earlier results where similar effects were presented in the context of topological superconductivity in a homogeneous wire, and attributed to more exotic time-reversal breaking mechanisms [Nat. Phys. 17, 43 (2020)]. In contrast, our stray field profiles are not designed to create Majorana modes, and our data are compatible with a straightforward interpretation in terms of trivial states in quantum dots. At the same time, the use of micromagnets in hybrid superconductor-semiconductor devices shows promise for future experiments on topological superconductivity.

Jiang, Luyao [Univ. of Pittsburgh, PA (United Stat↗

Matching Curved Lattices to Anisotropic Tangent Planes

Radial quantization would be the ideal formalism for studying strongly-coupled near-conformal quantum field theories but it requires the ability to perform lattice calculations on static, curved manifolds, specifically a very long cylinder whose cross section is a sphere. Smoothly discretizing the surface of a sphere requires a graph with unequal edge lengths. The geometry of such graphs is well understood since 1961 using Regge Calculus. But, lattice quantum field theories are defined in terms of couplings which appear in the action rather than edge lengths and so the relationship between couplings and lengths must be determined dynamically. A simple example is computing the ratio of spatial to temporal lattice spacings in anisotropic lattice QCD. I will discuss our conjecture that computing anisotropic lattice spacing ratios on affine transformations of regular flat lattices is sufficient to determine coupling assignments on curved lattices.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Astronomical Spectroscopy with Skipper CCDs: First Results from a Skipper CCD Focal Plane Prototype at SIFS

We present the first on-sky results from an ultra-low-readout-noise Skipper CCD focal plane prototype for the SOAR Integral Field Spectrograph (SIFS). The Skipper CCD focal plane consists of four 6k $\times$ 1k, 15 $\mu$m pixel, fully-depleted, $p$-channel devices that have been thinned to $\sim$250 $\mu$m, backside processed, and treated with an anti-reflective coating. These Skipper CCDs were configured for astronomical spectroscopy, i.e., a single-sample readout noise $< 4.3$\,e$^-$\,rms/pixel, the ability to achieve multi-sample readout noise $\ll 1\,$e$^-$\,rms/pix, full-well capacities $\sim$40,000--65,000 e$^-$, low dark current and charge transfer inefficiency ($\sim 2$ $\times$ $10^{-4}$ e$^-$/pixel/s and 3.44 $\times$ $10^{-7}$, respectively), and an absolute quantum efficiency of $\gtrsim 80\%$ between 450\,nm and 980\,nm ($\gtrsim 90\%$ between 600\,nm and 900\,nm). We optimize the readout sequence timing to achieve sub-electron noise ($\sim 0.5$\,e$^-$\,rms/pix) and photon-counting ($\sim 0.22$\,e$^-$\,rms/pix) readout noise over an area of 2k $\times$ 4k and 110 $\times$ 4k pixels, respectively in a readout time of $\lesssim 17$min. We observed two Lyman-$\alpha$ emitting quasars (HB89\,1159$+$123 and QSO\,J1621$–$0042) at redshift $z \sim 3.5$, two moderate redshift galaxy clusters (CL\,J1001$+$0220 and SPT-CL\,J2040$-$4451), an emission line galaxy at $z = 0.3239$, a candidate member star of the Bo\"{o}tes~II ultra-faint dwarf galaxy, and five CALSPEC spectrophotometric standard stars (HD074000, HD60753, HD106252, HD101452, HD200654). We present charge-quantized, photon-counting observation of the quasar HB89\,1159$+$123 and show the detector sensitivity increase for faint spectral features. We demonstrate signal-to-noise performance improvement for SIFS for observations in the low-background, readout-noise-dominated regime. We outline preliminary scientific studies that will leverage the SIFS-Skipper CCD data to derive new cosmological measurements in the context of dark matter and galaxy evolution.

79 ASTRONOMY AND ASTROPHYSICS↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Regime Characterization of Offshore Wind Resource Using Unsupervised Learning

Predictability of wind resource conditions is critical for offshore wind design and operations. While many studies of extreme wind conditions focus on specific events such as low-level jets or ramps, these rely on threshold definitions that limit generality. Here we present a data-driven framework that combines principal component analysis (PCA), self-organizing maps (SOM), and k-means clustering to classify wind resource conditions as typical and anomalous from climatological data. Anomalies are defined not by fixed thresholds but by flagging samples located far from SOM node centers inside the baseline SOM structure. This reframes extremes as rare ebents and hence, likely difficult to anticipate by numerical weather prediction models. We applied this approach to 23 years (2000–2022) of hourly profiles from the NOW-23 hindcast model at the Humboldt Wind Energy Area. Classification is conducted on a feature space consisting of 10 m wind speed and direction, bulk shear and veer across 30–270 m, and a low-level jet index. Dimensionality reduction is achieved through PC. A 2 × 3 OM lattice trained on the PCA vectors identified six baseline regimes spanning weak to strong flow states. High quantization-error profiles are identified and re-clustered into four anomalous regimes. The baseline regimes exhibited clear seasonal and diurnal cycles. Meanwhile, the anomalous regimes represented <10 % of all hours but showed distinct combinations of speed, shear, and veer, when compared to the baseline regimes. Anomalous regimes are typically short-lived (~few hours), yet their transitions can lead to hub-height wind changes of −18 to +9 m s -1 . For a representative 15 MW turbine, these shifts imply rapid swings in capacity factor from near-full output to negligible generation. Validation with lidar buoy data showed 51% agreement in SOM labels across ~6,000 overlapping hours, with most mismatches confined to adjacent speed classes. HRRR comparisons further revealed that anomalous regimes were disproportionately associated with forecast biases exceeding 5 m s -1 . Together, these results reframe extremes in offshore wind from absolute maxima or minima to weather states that are difficult to anticipate from models.

17 WIND ENERGY↗

Scalable workflow for evaluating and optimizing large language models

This work describes the improved workflow for evaluating open-source large language models (LLMs) for trustworthiness. The workflow facilitates the acquisition of LLMs, the generation of LLM responses, and the evaluation of the responses for their trustworthiness. As a use case, the workflow is employed to evaluate dense, quantized, and pruned Meta Llama3.1 LLMs for their truthfulness. The outcome of the project could set the stage for understanding and developing trustworthy models in the future projects.

97 MATHEMATICS AND COMPUTING↗

Energy-efficient, Large-scale Molecular Dynamics Simulations via Hardware- and Algorithm-level Optimization

This work aims to develop a framework for energy-efficient computing that will enable molecular dynamics (MD) simulations of large-scale phenomena with atomic precision and simultaneously remove computational bottlenecks limiting the speed of MD simulations. We seek to implement such an approach through the development of surrogate models for the interatomic force calculation combined with the use of mixed numerical precision formats. For a model system of neutral atoms (only pairwise interactions), significant force calculation efficiency improvements were achieved, without detrimental effects on atomic structures or average energies, using single precision, by developing a surrogate model (deep neural network), and by quantizing this surrogate model. For a model system of charged atoms, the reciprocal-space calculation of electrostatic interactions was identified as the main bottleneck, and the development of a surrogate model should be pursued to achieve an estimated one-order-of-magnitude additional speedup.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Electron-Proton Scattering Event Generation using Structured Tokenization

Recent work such as Omnijet-$\alpha$ has demonstrated that effective tokenization combined with transformer-based architectures can produce effective foundation models for jet physics. While tokenization may help models capture generalizable event characteristics, it also introduces discretization errors that may compromise the precision required for downstream physics analyses. As the number and complexity of the particle features grow, these errors are likely to grow proportionally. In this study, we investigate new tokenization strategies to improve the application of generative transformer models to \textsc{Pythia8} simulations of electron-proton scattering at the Electron-Ion Collider. Specifically, we propose a feature-based structured tokenization approach that utilizes multiple tokens per particle, improving expressivity, while reducing the total number of unique tokens needed. We evaluate this method against grid-based binning, K-means clustering, and vector-quantized variational auto-encoders on the event simulations. Our results show that feature-based structured tokenization reduces discretization error, leading to more accurate generative modeling of particle-level events.

Goldenberg, Steven [Thomas Jefferson National Acce↗

UNLOQ: UNderstanding coherence in Light-matter interfaces for Quantum Science (Final Technical Report)

The general goal of this project is to prepare next-generation quantum systems for novel quantum information science applications. Decades of research on quantum optics have provided revolutionary systems for manipulating atomic and photonic quantum coherence in transformative ways. Our hypothesis is that the next-generation quantum systems will come from nano-molecular quantum optics. Specifically, we are interested in coupling the electronic states of molecules or nanoparticles to the quantized radiation field inside an optical cavity to create a set of new photon-matter hybrid excitations, called polaritons. As opposed to atoms, the vibrational modes of molecules and nanoparticles provide new degrees of freedom to mediate the quantum transduction between electronic and photonic states, offering new ways to tune and ultimately control the quantum coherence of the integrated system. In this project we are interested in designing, fabricating, and characterizing the polaritons that arise from coupling CdSe nanoplatelets (NPLs) to a Fabry-Pérot optical cavity. Specific attention was paid to parameters of the system (cavity mode volume, quality factor, etc.) that would maximize the collective coupling strength. Through a combined theoretical and experimental approach, we were able to make important contributions to understanding NPL exciton-polariton photophysics including how to use cavity loss as a tunable parameter in order to manipulate the populations of the upper and lower polariton states. Finally, using state-of-the-art theoretical methods, we were able to show how the collective coupling of many molecular excitons in a cavity can protect polariton coherence from vibrationally-induced decoherence. In particular, for NPL-cavity polaritonic systems, the quantum coherence can be extended by over an order of magnitude at room temperature.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Low-latency Jet Tagging for HL-LHC Using Transformer Architectures

Transformers are the state-of-the-art model architectures and widely used in application areas of machine learning. However the performance of such architectures is less well explored in the ultra-low latency domains where deployment on FPGAs or ASICs is required. Such domains include the trigger and data acquisition systems of the LHC experiments. We present a transformer-based algorithm for jet tagging built with the HGQ2 framework, which is able to produce a model with heterogeneous bitwidths for fast inference on FPGAs, as required in the trigger systems at the LHC experiments. The bitwidths are acquired during training by minimizing the total bit operations as an additional parameter. By allowing a bitwidth of zero, the model is pruned in-situ during training. Using this quantization-aware approach, our algorithm achieves state-of-the-art performance while also retaining permutation invariance which is a key property for particle physics applications. Due to the strength of transformers in representation learning, our work also serves as a stepping stone for the development of a larger foundation model for trigger applications.

Laatu, Lauri [Imperial Coll., London]↗