Comparative analysis of model compression techniques for achieving carbon efficient AI
Not provided.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not provided.
For a gas–solid interfacial system where chemical species undergo reversible adsorption, we develop a mesoscopic stochastic modeling method that simulates both gas-phase hydrodynamics and surface coverage dynamics by coupling the Langmuir adsorption model with compressible fluctuating hydrodynamics. To this end, we derive a thermodynamically consistent mass–energy update scheme that accounts for how the mass and energy variables in the gas and surface subsystems should be updated according to the changes in the number of molecules of each species in each subsystem due to adsorption and desorption events. By performing a stochastic analysis for the ideal Langmuir model and the full hydrodynamic system, we analytically confirm that our mass–energy update scheme captures thermodynamic equilibrium predicted by equilibrium statistical mechanics. We find that an internal energy correction term is needed, which is attributed to the difference in the mean kinetic energy of gas molecules colliding with the surface from that computed from the Maxwell–Boltzmann distribution. By performing an equilibrium simulation study for an ideal gas mixture of CO and Ar, with CO undergoing reversible adsorption, we validate our overall simulation method and implementation.
In this work, a near-wall model, which couples the inverse of a recently developed compressible velocity transformation (Griffin et al., Proc. Natl Acad. Sci., vol. 118, 2021, p. 34) and an algebraic temperature–velocity relation, is developed for high-speed turbulent boundary layers. As input, the model requires the mean flow state at one wall-normal height in the inner layer of the boundary layer and at the boundary-layer edge. As output, the model can predict mean temperature and velocity profiles across the entire inner layer, as well as the wall shear stress and heat flux. The model is tested in an a priori sense using a wide database of direct numerical simulation high-Mach-number turbulent channel flows, pipe flows and boundary layers (48 cases, with edge Mach numbers in the range 0.77–11, and semi-local friction Reynolds numbers in the range 170–5700). The present model is significantly more accurate than the classical ordinary differential equation (ODE) model for all cases tested. The model is deployed as a wall model for large-eddy simulations in channel flows with bulk Mach numbers in the range 0.7–4 and friction Reynolds numbers in the range 320–1800. When compared to the classical framework, in the a posteriori sense, the present method greatly improves the predicted heat flux, wall stress, and temperature and velocity profiles, especially in cases with strong heat transfer. In addition, the present model solves one ODE instead of two, and has a computational cost and implementation complexity similar to that of the commonly used ODE model.
Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.
Official code for the NeurIPS 2023 paper "Neural Image Compression: Generalization, Robustness, and Spectral Biases". This code can be used to: - Computing spectral distortion errors of images compressed with traditional codes or neural image compression models (as proposed in the aformentioned paper) - Visualize power spectral densities and fourier heatmaps (as proposed in the aformentioned paper) - Train and test a variety of neural image compression models
Network complexity and computational efficiency have become increasingly significant aspects of deep learning. Sparse deep learning addresses these challenges by recovering a sparse representation of the underlying target function by reducing heavily overparameterized deep neural networks. Specifically, deep neural architectures compressed via structured sparsity (e.g., node sparsity) provide low-latency inference, higher data throughput, and reduced energy consumption. In this article, we explore two well-established shrinkage techniques, Lasso and Horseshoe, for model compression in Bayesian neural networks (BNNs). To this end, we propose structurally sparse BNNs, which systematically prune excessive nodes with the following: 1) spike-and-slab group Lasso (SS-GL) and 2) SS group Horseshoe (SS-GHS) priors, and develop computationally tractable variational inference, including continuous relaxation of Bernoulli variables. We establish the contraction rates of the variational posterior of our proposed models as a function of the network topology, layerwise node cardinalities, and bounds on the network weights. Furthermore, we empirically demonstrate the competitive performance of our models compared with the baseline models in prediction accuracy, model compression, and inference latency.
Explore the source record for details and available documents.
Knowledge distillation is a form of model compression that allows artificial neural networks of different sizes to learn from one another. Its main application is the compactification of large deep neural networks to free up computational resources, in particular on edge devices. In this article, we consider proton-proton collisions at the High-Luminosity Large Hadron Collider (HL-LHC) and demonstrate a successful knowledge transfer from an event-level graph neural network (GNN) to a particle-level small deep neural network (DNN). Our algorithm, DistillNet, is a DNN that is trained to learn about the provenance of particles, as provided by the soft labels that are the GNN outputs, to predict whether or not a particle originates from the primary interaction vertex. The results indicate that for this problem, which is one of the main challenges at the HL-LHC, there is minimal loss during the transfer of knowledge to the small student network, while improving significantly the computational resource needs compared to the teacher. This is demonstrated for the distilled student network on a CPU, as well as for a quantized and pruned student network deployed on a field programmable gate array. Our study proves that knowledge transfer between networks of different complexity can be used for fast artificial intelligence (AI) in high-energy physics that improves the expressiveness of observables over non-AI-based reconstruction algorithms. Such an approach can become essential at the HL-LHC experiments, e.g. to comply with the resource budget of their trigger stages.
Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.
We present a new finite element framework for modeling compressible, turbulent multiphase flows with heat transfer. For two-fluid systems with a free surface, the Volume of Fluid (VOF) method is implemented without the need for interface reconstruction, while turbulence is resolved using a dynamic Vreman large eddy simulation (LES) model. Unlike most two-phase VOF studies, which neglect heat transfer, the present approach incorporates energy transport equations within the VOF formulation to account for heat exchange, an effect particularly important in turbulent flows. Conjugate heat transfer is often challenging in finite volume methods, which require explicit specification of heat fluxes at the solid–fluid interface, limiting accuracy and predictive capability. By contrast, the finite element formulation does not require heat flux inputs, allowing more accurate and robust simulation of heat transfer between solids and fluids. The method is demonstrated through three representative cases. First, a two-fluid instability with a single-mode perturbation is simulated and validated against analytical growth rates. Second, conjugate heat transfer is examined in a high-temperature flow over a cold metal cylinder, with validation performed both quantitatively—via pressure coefficient comparisons with experimental data—and qualitatively using vector field topology. Finally, compressible spray injection and breakup are modeled, demonstrating the ability of the framework to capture interfacial dynamics and atomization under turbulent, high-speed conditions. In the compressible spray injection and breakup case, the results indicate that the finite element formulation achieved higher predictive accuracy and robustness than the finite-volume method. With the same mesh resolution, the FEM reduced the root mean square error (RMSE) and mean absolute percentage error (MAPE) from 6.96 mm and 26.0% (for the FVM) to 4.85 mm and 12.7%, respectively, demonstrating improved accuracy and robustness in capturing interfacial dynamics and heat transfer. The study also introduced vector field topology to visualize and interpret coherent flow structures and instabilities, offering insights beyond conventional scalar-field analyses.
Abstract Compact symbolic expressions have been shown to be more efficient than neural network (NN) models in terms of resource consumption and inference speed when implemented on custom hardware such as field-programmable gate arrays (FPGAs), while maintaining comparable accuracy (Tsoi et al 2024 EPJ Web Conf. 295 09036). These capabilities are highly valuable in environments with stringent computational resource constraints, such as high-energy physics experiments at the CERN Large Hadron Collider. However, finding compact expressions for high-dimensional datasets remains challenging due to the inherent limitations of genetic programming (GP), the search algorithm of most symbolic regression (SR) methods. Contrary to GP, the NN approach to SR offers scalability to high-dimensional inputs and leverages gradient methods for faster equation searching. Common ways of constraining expression complexity often involve multistage pruning with fine-tuning, which can result in significant performance loss. In this work, we propose S y m b o l N e t , a NN approach to SR specifically designed as a model compression technique, aimed at enabling low-latency inference for high-dimensional inputs on custom hardware such as FPGAs. This framework allows dynamic pruning of model weights, input features, and mathematical operators in a single training process, where both training loss and expression complexity are optimized simultaneously. We introduce a sparsity regularization term for each pruning type, which can adaptively adjust its strength, leading to convergence at a target sparsity ratio. Unlike most existing SR methods that struggle with datasets containing more than O ( 10 ) inputs, we demonstrate the effectiveness of our model on the LHC jet tagging task (16 inputs), MNIST (784 inputs), and SVHN (3072 inputs).
Large-scale classical molecular dynamics (CMD) simulations naturally include the microscopic physics necessary for atomistic modeling of shock release at the ablator-fuel interface in an inertial confinement fusion (ICF) capsule. Here, the multi-megabar shocks utilized in ICF experiments can drive the deuterium fuel from ambient to electron volt temperatures (T) and multi-fold compression. Modeling interatomic interactions over such an extreme range of conditions is challenging for empirical bond order potentials. We generate a pair potential for deuterium with explicit temperature and mass density dependence from ab initio density functional theory molecular dynamics using the iterative Boltzmann inversion method. This potential accurately reproduces the radial distribution functions and pressures from DFT in CMD equilibrium simulations across a wide range of thermodynamic conditions, yet fails to return the expected Hugoniot relations when used in direct CMD shock simulations.
DeePMD-kit is a powerful open-source software package that facilitates molecular dynamics simulations using machine learning potentials known as Deep Potential (DP) models. This package, which was released in 2017, has been widely used in the fields of physics, chemistry, biology, and material science for studying atomistic systems. The current version of DeePMD-kit offers numerous advanced features, such as DeepPot-SE, attention-based and hybrid descriptors, the ability to fit tensile properties, type embedding, model deviation, DP-range correction, DP long range, graphics processing unit support for customized operators, model compression, non-von Neumann molecular dynamics, and improved usability, including documentation, compiled binary packages, graphical user interfaces, and application programming interfaces. This article presents an overview of the current major version of the DeePMD-kit package, highlighting its features and technical details. Additionally, this article presents a comprehensive procedure for conducting molecular dynamics as a representative application, benchmarks the accuracy and efficiency of different models, and discusses ongoing developments.
Steam piping networks are essential for optimizing performance in industrial processes and district heating systems. However, dynamic models that balance thermo-hydraulic accuracy with computational efficiency remain limited. In response, this paper presents a new discretized steam pipe model based on the plug flow approach, capturing key thermo-hydraulic behaviors while simplifying steam phase change processes. Implemented in Modelica, the model accurately calculates temperature and pressure distributions along steam pipelines. To improve computational efficiency for district-scale simulations, five model simplifications are introduced: lumped thermo-hydraulic functions, empirical correlations, fluid state approximations, steady-state dynamics and inclusion of flow derivatives. These simplified models achieve 85%-98% accuracy in predicting pressure drop and condensation losses, including dynamic condensate behavior during pipe warm-up—a factor often overlooked in existing models. The models support diverse network configurations, scaling effectively to systems with multiple distribution pipes and connected building loads. Discrete models provide detailed insights but exhibit a cubic increase in simulation time as the network scales by N connected building O(N 2.42 ). In contrast, lumped models simulate 10–28 times faster than discrete, offering quadratic scaling of simulation time O(N 1.73 ). However, they still require 6 times more computation time than a lossless network, highlighting the inherent computational challenges of modeling compressible fluid flow. In conclusion, the steady-state lumped variant, with its near-linear scalability in computational time O(N 1.01 ), emerges as an efficient solution for preliminary design evaluations and extensive parametric studies.
This codebase uses machine learning to train a collection of 3D Gaussian distributions to approximate scientific volume data. Because this collection uses less memory than the original dataset, it can be used as a compressed model of the original data for applications such as visualization.
X-ray crystallography reconstruction, which transforms discrete X-ray diffraction patterns into three-dimensional molecular structures, relies critically on accurate Bragg peak finding for structure determination. As X-ray free electron laser (XFEL) facilities advance toward MHz data rates (1 million images per second), traditional peak finding algorithms that require manual parameter tuning or exhaustive grid searches across multiple experiments become increasingly impractical. While deep learning approaches offer promising solutions, their deployment in high-throughput environments presents significant challenges in automated dataset labeling, model scalability, edge deployment efficiency, and distributed inference capabilities. We present an end-to-end deep learning pipeline with three key components: (1) a data engine that combines traditional algorithms with our peak matching algorithm to generate high-quality training data at scale, (2) a modular architecture that scales from a few million to hundreds of million parameters, enabling us to train large expert-level models offline while deploying smaller, distilled models at the edge, and (3) a decoupled producer-consumer architecture that separates specialized data source layer from model inference, enabling flexible deployment across diverse computing environments. Using this integrated approach, our pipeline achieves accuracy comparable to traditional methods tuned by human experts while eliminating the need for experiment-specific parameter tuning. Although current throughput requires optimization for MHz facilities, our system's scalable architecture and demonstrated model compression capabilities provide a foundation for future high-throughput XFEL deployments.
The Single Primary Heat Extraction and Rejection Emulator (SPHERE) facility at Idaho Na- tional Laboratory was recently utilized to generate data for the startup and steady operation of a high-performance, sodium heat pipe over the course of 1000 hours, as a test of detrimental, long-term effects of heat pipe operation. The setup consists of a single, sodium heat pipe enclosed in a stainless-steel vacuum chamber, heated radiatively via a cylindrical ceramic fiber heater configuration and cooled via a water-cooled calorimeter. Measurements include temperatures at several axial locations along the outer surface of the heat pipe, the power provided to the heaters, and the heat removal rate of the calorimeter. In this work, this data is utilized to validate heat pipe models in the heat pipe application Sockeye, which is based upon the Multiphysics Object-Oriented Simulation Environment (MOOSE) framework. Sockeye provides various heat pipe models at an engineering scale appropriate for the multiphysics simulation of microreactors, which may feature several hundred heat pipes. This work details models of this experiment in SPHERE using various heat pipe models with Sockeye, including heat conduction-based models and compressible flow models of the heat pipe interior. These models are compared to the experimental data to assess the accuracy of several aspects of heat pipe modeling, including frozen startup, the effect of non-condensable gases, and the coupling of the heat pipe to its environment.