Search NASA⌕ Search

SEARCH · Search NASA

Results for “data compaction and compression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Progressive Tree-Based Compression of Large-Scale Particle Data

Scientific simulations and observations using particles have been creating large datasets that require effective and efficient data reduction to store, transfer, and analyze. However, current approaches either compress only small data well while being inefficient for large data, or handle large data but with insufficient compression. Toward effective and scalable compression/decompression of particle positions, we introduce new kinds of particle hierarchies and corresponding traversal orders that quickly reduce reconstruction error while being fast and low in memory footprint. Our solution to compression of large-scale particle data is a flexible block-based hierarchy that supports progressive, random-access, and error-driven decoding, where error estimation heuristics can be supplied by the user. For low-level node encoding, we introduce new schemes that effectively compress both uniform and densely structured particle distributions. Our proposed methods thus target all three phases of a tree-based particle compression pipeline, namely tree construction, tree traversal, and node encoding. In conclusion, the improved efficacy and flexibility of these methods over existing compressors are demonstrated through extensive experimentation, using a wide range of scientific particle datasets.

97 MATHEMATICS AND COMPUTING↗

A 28 nm multiply-accumulate ASIC architecture for on-chip data compression in MHz frame rate X-ray and electron pixel detectors

Modern X-ray detector systems urgently require compact, efficient, and fast data compression schemes to handle the transmission of big data from pixel arrays, enabling frame rates in the MHz regime. Here, in this work, a data compression ASIC that implements a streaming fixed-length lossy compression scheme is introduced and analyzed, proving the feasibility and benefits of on-chip compression. The compression scheme utilizes a vector matrix product logic, which performs a number of floating-point multiplications, additions, and accumulations. The logic is verified, synthesized, and shown to fit in the area resource available for the X-ray detector under study, which comprises 192 × 168 pixels each of 12-bit width, and having a total area of 20 mm× 20 mm, about 2 mm× 20 mm of which are available for the digital logic. Several system architectures, precisions, and compression ratios ranging from 100 to 250 were analyzed to pave the way for on-chip fixed-length compression (e.g., principal component analysis, singular value decomposition) and data reduction (e.g., azimuthal integration) for X-ray and electron detectors.

Data compression↗

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗

NeRVI: Compressive neural representation of visualization images for communicating volume visualization results

We present NeRVI, a new deep-learning approach that compresses a large collection of visualization images generated from time-varying data for communicating volume visualization results. Based on an image-based implicit neural representation, our approach represents tens of thousands of high-resolution rendering images parametrized by different parameters via a hybrid model of multilayer perceptrons and convolutional neural networks. Here, our model predicts images and corresponding masks, and the masks are utilized for loss computation and network training to capture fine structural details and small components. In conjunction with model quantization and weight encoding, NeRVI yields highly compact compressive neural representations while preserving the image fidelity well. We demonstrate the effectiveness of NeRVI with isosurface rendering and direct volume rendering images generated from multiple data sets and compare NeRVI with other state-of-the-art deep learning-based (InSituNet, SIREN, NeRF, and NeRV) methods. Quantitative and qualitative results show that NeRVI provides an alternative solution that augments domain scientists' ability to manage, represent, and communicate scientific visualization output.

97 MATHEMATICS AND COMPUTING↗

Toward first principles-based simulations of dense hydrogen

Accurate knowledge of the properties of hydrogen at high compression is crucial for astrophysics (e.g., planetary and stellar interiors, brown dwarfs, atmosphere of compact stars) and laboratory experiments, including inertial confinement fusion. There exists experimental data for the equation of state, conductivity, and Thomson scattering spectra. However, the analysis of the measurements at extreme pressures and temperatures typically involves additional model assumptions, which makes it difficult to assess the accuracy of the experimental data rigorously. On the other hand, theory and modeling have produced extensive collections of data. They originate from a very large variety of models and simulations including path integral Monte Carlo (PIMC) simulations, density functional theory (DFT), chemical models, machine-learned models, and combinations thereof. At the same time, each of these methods has fundamental limitations (fermion sign problem in PIMC, approximate exchange–correlation functionals of DFT, inconsistent interaction energy contributions in chemical models, etc.), so for some parameter ranges accurate predictions are difficult. Recently, a number of breakthroughs in first principles PIMC as well as in DFT simulations were achieved which are discussed in this review. Here we use these results to benchmark different simulation methods. We present an update of the hydrogen phase diagram at high pressures, the expected phase transitions, and thermodynamic properties including the equation of state and momentum distribution. Furthermore, we discuss available dynamic results for warm dense hydrogen, including the conductivity, dynamic structure factor, plasmon dispersion, imaginary-time structure, and density response functions. We conclude by outlining strategies to combine different simulations to achieve accurate theoretical predictions that are based on first principles.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Concurrent measurement of strain and chemical reaction rates in a calcite grain pack undergoing pressure solution: Evidence for surface-reaction controlled dissolution

Pressure solution is inferred to be a significant contributor to sediment compaction and lithification, especially in carbonate sediments. For a sediment deforming primarily by pressure solution, the compaction rate should be directly related to the rate of calcite dissolution, transport along grain contacts, and calcite reprecipitation. Previous experimental work has shown that there is evidence that deformation in wet calcite grain packs is consistent with control by pressure solution, but considerable ambiguity remains regarding the rate limiting mechanism. We present the results of laboratory compaction experiments designed to directly measure calcite dissolution and precipitation rates (recrystallization rates) concurrently with strain rate to test whether measured rates are consistent with predicted rates both in absolute magnitude and time evolution. Recrystallization rates are measured using trace element chemistry (Sr/Ca, Mg/Ca) and isotopes (87Sr/86Sr) of fluids flowing slowly through a compacting grain pack as it is being triaxially compressed. Imaging techniques are used to characterize the grain contacts and strain effects in the post-experiment grain pack. Our data show that calcite recrystallization rates calculated from all three geochemical parameters are in approximate agreement and that the rates closely track strain rate. The geochemically inferred rates are close to predicted rates in absolute magnitude. Uncertainty in grain contact dimensions makes distinguishing between surface reaction control and diffusion control difficult. Measured reaction rates decrease faster than predicted from standard pressure solution creep flow laws. This inconsistency may indicate that calcite dissolution rates at grain contacts are more complex, and more time-dependent, than suggested by geometric models designed to predict grain contact stresses.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Discrete element model for powder grain interactions under high compressive stress

A reduced order, nonlocal model is proposed for the contact force between initially spherical particles under compression. The model in effect provides the normal component of the interaction force between elements in the discrete element method (DEM). It is applicable to high relative density and large stress in powder compaction. It takes into account the mutual interaction between multiple points of contact, in contrast to the usual assumption in DEM of pair interactions. The mathematical form of the model is derived from a variational formulation that leads to the momentum balance for the forces on each grain. The model is calibrated mainly using detailed three dimensional peridynamic simulations of single grains under compressive loading by rigid plates that move radially with prescribed velocity. This calibration takes into account the large deformation and fracture of the grains. The interaction model also includes terms for the unloading behavior and adhesion. Finally, as validation, the model is applied to test data on the compaction of microcrystalline cellulose bulk powder.

36 MATERIALS SCIENCE↗

Glassy carbon formation from pyrolysis of polymeric coatings on fiber-optic sensors

Deploying fiber-optic sensors in nuclear reactors requires a detailed understanding of radiation effects on the fiber materials and the transmitted signals. Previous work has shown large wavelength shifts in the reflected spectra obtained from polymer-coated fiber-optic temperature sensors exposed to high neutron fluences. The sensor drift resulting from these wavelength shifts cannot be explained by radiation effects on fused silica glass. These shifts are hypothesized to be caused by the conversion of the polymeric fiber coating to a glassy carbon via radiolysis and/or pyrolysis and subsequent radiation-induced compaction. Here, thermal degradation of these polymeric coatings was studied to provide insight into the potential origins of the sensor drift phenomenon. Acrylate- and polyimide-coated fibers were heated under various temperatures (250–1300 °C) and environments (oxidative and inert), and the resulting coating products were characterized via mass-loss data, scanning electron microscope imaging, and Raman spectroscopy. Results suggest that the polymer decomposition product of both coating types, at least under inert conditions, is indeed a glassy carbon. Analytical models that account for radiation-induced glassy carbon coating compaction show significant compressive fiber strains and predicted wavelength shifts that agree well with experimental measurements, providing additional evidence that supports the hypothesized origins of the sensor drift.

36 MATERIALS SCIENCE↗

Integration of RNTuple in ATLAS Athena

After using ROOT’s TTree I/O subsystem for over two decades and storing more than an exabyte of compressed High Energy Physics (HEP) data, advances in technology have motivated a complete redesign, RNTuple, which breaks backward-compatibility to take better advantage of these storage options. The RNTuple I/O subsystem has been designed to address performance bottlenecks and other shortcomings of TTree. Specifically, RNTuple comes with an updated, more compact binary data format that can be stored both in ROOT files and natively in object stores. It is designed for modern storage hardware (e.g. high-throughput low-latency NVMe SSDs), and provides robust and easy to use interfaces. The binary format of RNTuple is scheduled to become production grade in 2024, and recently has become mature enough to start exploring the integration into software used by HEP experiments. In this contribution, we discuss the developments to support the features as required by the ATLAS analysis Event Data Model (EDM) in RNTuple, which will enable its integration into the Athena software framework. With these developments in place, we evaluate the performance of the current most recent versions of RNTuple-based ATLAS data sets and compare this to that of TTree.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

CMS High Granularity Calorimeter ECON-D ASIC overview and radiation testing results

The Compact Muon Solenoid (CMS) Experiment’s High Granularity Calorimeter (HGCAL) upgrade replaces the CMS electromagnetic and hadronic endcap calorimeters in preparation for the high-rate and high-radiation environment of the High Luminosity LHC. To effectively use the over 6 million channels of this highly-segmented “imaging” calorimeter, CMS is pioneering very front-end data compression with the Endcap Concentrator (ECON) ASICs – the ECON-T for the trigger path and the ECON-D for the data path. These 65 nm CMOS ASICs are radiation tolerant (200 Mrad) and low-power (< 2.5 mW/channel). In June 2023, we received the first full-functionality prototype of the data path concentrator ASIC, the ECON-D-P1. This talk will present an overview of the ECON-D-P1, summarize functionality and system testing, and present results from both Total Ionizing Dose (TID) and Single Event Effect (SEE) testing campaigns completed in summer 2023, validating the ECON-D radiation tolerant performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Dynamic Low-Rank Training with Spectral Regularization: Achieving Robustness in Compressed Representations

Deployment of neural networks on resource-constrained devices demands models that are both compact and robust to adversarial inputs. However, compression and adversarial robustness often conflict. In this work, we introduce a dynamical low-rank training scheme enhanced with a novel spectral regularizer that controls the condition number of the low-rank core in each layer. This approach mitigates the sensitivity of compressed models to adversarial perturbations without sacrificing clean accuracy. The method is model- and data-agnostic, computationally efficient, and supports rank adaptivity to automatically compress the network at hand. Extensive experiments across standard architectures, datasets, and adversarial attacks show the regularized networks can achieve over 94\% compression while recovering or improving adversarial accuracy relative to uncompressed baselines.

Schotthoefer, Steffen [ORNL] (ORCID:00000002156965↗

Dynamical Low-Rank Compression of Neural Networks with Robustness under Adversarial Attacks

Deployment of neural networks on resource-constrained devices demands models that are both compact and robust to adversarial inputs. However, compression and adversarial robustness often conflict. In this work, we introduce a dynamical low-rank training scheme enhanced with a novel spectral regularizer that controls the condition number of the low-rank core in each layer. This approach mitigates the sensitivity of compressed models to adversarial perturbations without sacrificing clean accuracy. The method is model- and data-agnostic, computationally efficient, and supports rank adaptivity to automatically compress the network at hand. Extensive experiments across standard architectures, datasets, and adversarial attacks show the regularized networks can achieve over 94 compression while recovering or improving adversarial accuracy relative to uncompressed baselines.

Schotthoefer, Steffen [ORNL] (ORCID:00000002156965↗

Inverse design of cellular structures with the targeted nonlinear mechanical response

Advanced additive manufacturing capabilities have enabled a transformational ability to create sophisticated cellular structures using diverse materials. By altering the topology of the unit cell, the mechanical behavior, such as the stress-strain response during compression, can be modulated. Nevertheless, identifying a printable topology within an enormous design space that would precisely deliver the targeted nonlinear material response is challenging. We propose a data-driven generative framework based on a conditional variational autoencoder (cVAE) architecture that can inverse design the cellular structure based on the intended nonlinear stress-strain response. Trained on a dataset of structure-property pairs, the cVAE learns a compact and expressive latent space that enables efficient mapping from targets to feasible geometries. Two inference modes are explored: (1) decoder-only generation, which enables the exploration of diverse designs conditioned solely on the desired mechanical response, and (2) encoder-decoder generation, which further allows for the incorporation of desired topologies, ensuring the generated structure conforms to both mechanical properties and to desired-topology constraints. The results demonstrate that the model can generate structurally plausible and mechanically accurate designs, with the predicted stress-strain curves closely matching the targets. Even under joint conditioning, the model effectively balances geometric fidelity and functional performance.

36 MATERIALS SCIENCE↗

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

Lifting MGARD: Construction of (pre)wavelets on the interval using polynomial predictors of arbitrary order

MGARD (MultiGrid Adaptive Reduction of Data) is an algorithm for compressing and refactoring scientific data, based on the theory of multigrid methods. The core algorithm is built around stable multilevel decompositions of conforming piecewise linear $C^0$ finite element spaces, enabling accurate error control in various norms and derived quantities of interest. In this work, we extend this construction to arbitrary order Lagrange finite elements $\mathbb{Q}_p$, $p \geq 0$, and propose a reformulation of the algorithm as a lifting scheme with polynomial predictors of arbitrary order. Additionally, a new formulation using a compactly supported wavelet basis is discussed, and an explicit construction of the proposed wavelet transform for uniform dyadic grids is described.

Reshniak, Viktor [Oak Ridge National Laboratory (O↗

Genetic algorithm optimization of a chemical kinetic mechanism for propane at engine relevant conditions

Propane has demonstrated significant potential for reductions in greenhouse gas and pollutant emissions in medium- and heavy-duty engine applications, but further improvements require accurate, compact, and scalable chemical kinetic mechanisms to design the next generation of propane fueled engines, particularly at the boosted operating conditions necessary to meet the power density demand of medium- and heavy-duty applications. In this work, six key chemical reactions were identified in a reduced mechanism with 70 species and 352 reactions through a sensitivity analysis performed at conditions typical of thermodynamic trajectories observed in a high compression ratio, long stroke engine operated on propane from throttled to boosted operating conditions. While the original mechanism was validated against rapid compression machine (RCM) data, it was found to overpredict experimental autoignition tendencies in 2-zone, 0-D SI engine simulations performed in Chemkin Pro. Subsequently, a genetic algorithm approach was used to optimize the six reaction rate parameters within established uncertainty bounds by performing RCM simulations and comparing to two independent sets of literature ignition delay times for propane, thus generating two new kinetic mechanisms. The first optimization achieved a mean absolute percent error (MPE) reduction in 2nd stage ignition delay of 61.4% in seven generations, while the second optimization utilized a newer experimental RCM dataset, and achieved MPE reduction of 56.7% in seven generations, and further marginal improvement to 57.8% reduction in 34 generations. Finally, the two mechanisms were then evaluated again in the 2-zone 0-D SI engine model in Chemkin Pro comparing typical mean and knocking cycle trajectories, and it was found that the second optimized mechanism provided better prediction of knock onset at the representative conditions evaluated in this work, particularly for higher load operating conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

F-Hash: Feature-Based Hash Design for Time-Varying Volume Visualization via Multi-Resolution Tesseract Encoding

Interactive time-varying volume visualization is challenging due to its complex spatiotemporal features and sheer size of the dataset. Recent works transform the original discrete time-varying volumetric data into continuous Implicit Neural Representations (INR) to address the issues of compression, rendering, and super-resolution in both spatial and temporal domains. However, training the INR takes a long time to converge, especially when handling large-scale time-varying volumetric datasets. In this work, we proposed F-Hash, a novel feature-based multi-resolution Tesseract encoding architecture to greatly enhance the convergence speed compared with existing input encoding methods for modeling time-varying volumetric data. The proposed design incorporates multi-level collision-free hash functions that map dynamic 4D multi-resolution embedding grids without bucket waste, achieving high encoding capacity with compact encoding parameters. Our encoding method is agnostic to time-varying feature detection methods, making it a unified encoding solution for feature tracking and evolution visualization. Experiments show the F-Hash achieves state-of-the-art convergence speed in training various time-varying volumetric datasets for diverse features. We also proposed an adaptive ray marching algorithm to optimize the sample streaming for faster rendering of the time-varying neural representation.

deep learning↗

The Role of Unit-Cell Topology in Modulating the Compaction Response of Additively Manufactured Cellular Materials using Simulations and Validation Experiments

Additive manufacturing has enabled a transformational ability to create cellular structures (or foams) with tailored topology. Compared to their monolithic polymer counterparts, cellular structures are potentially suitable for systems requiring materials with high specific energy-absorbing capability to provide enhanced damping. In this work, we demonstrate the utility of controlling unit-cell topology with the intent of obtaining a desired stress–strain response and energy density. Using mesoscale simulations that resolve the unit-cell sub-structures, we validate the role of unit-cell topology in selectively activating a buckling mode and thereby modulating the characteristic stress–strain response. Simulations incorporate a linear viscoelastic constitutive model and a hyperelastic model for simulating large deformation of the polymer under both tension and compression. Simulated results for nine different cellular structures are compared with experimental data to gain insights into three different modes of buckling and the corresponding stress–strain response.

36 MATERIALS SCIENCE↗