Search NASASearch

SEARCH · Search NASA

Results for “data compression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Pressure-temperature equation of state of Al 2 ⁢O 3 up to 14 Mbar and 40 kK

Sapphire (Al 2 ⁢O 3 ), known for its remarkable incompressibility at ambient conditions, plays a pivotal role in both static and dynamic compression research. Accurately characterizing its equation of state (EoS) is essential for these applications. Here, we present a complete Hugoniot of Al 2 ⁢ O 3 as locus of experimentally assessed, high-precision, pressure, density and temperature states up to 14 Mbar and 43 kK. The Hugoniot is established with single shock experiments using magnetically launched hyper velocity flyers on the Z Accelerator at Sandia National Laboratories. We explore principal Hugoniot states at very high shock 𝑇 and 𝑝 in the solid phase, tracking the solid-liquid boundary and culminating at 2.4-fold compression, where data provides a direct constraint on the liquid phase. Corresponding shock release data probe thermodynamic states complementary to the Hugoniot and place additional constraints on tabular EoS models. Our findings indicate a significant deviation from existing tabular EoS models for Al 2 ⁢ O 3 dictating a comprehensive overhaul. We develop two advanced EoSs for Al 2 ⁢ O 3 the SESAME 97412 model, featuring an extensive phase diagram that includes three solid phases and the liquid phase, and the updated LEOS 2200m2 model. EoS development is assisted with Quantum Molecular Dynamics simulations. Our experimental data allows for stringent testing of our EoSs. Both models accurately capture the Hugoniot of Al 2 ⁢O 3 up to the highest pressures and temperatures. Rigorous experimental determination of extreme pressures and temperatures, paired with sophisticated models, advances the frontier of EoS development beyond 1 terapascal.

Kalita, Patricia [Sandia National Laboratories (SN

NLR Data Processing Pipeline for MADIS [SWR-26-050]

The NLR Data Processing Pipeline for MADIS software package is for downloading, processing, and performing QA/QC on MADIS data. Designed to handle the following steps: 1) Download all MADIS data as compressed netcdf files for a given time period. 2) Unpack netcdf files into timeseries csvs for each coordinate within the given bounding box. 3) Process the csvs to filter according to quality control checks and convert variables to correct units. 4) Write processed csvs to a single nc file.

Benton, Brandon [National Laboratory of the Rockie

Frameworks, Algorithms, and Scalable Technologies for Mathematics (FASTMath) SciDAC Institute

As computational models scale to larger computers, the rate at which they produce data has far outstripped the same computers ability to write that data and further the file systems ability to store that data. Almost all of the SciDAC applications, but especially those related to fusion solve very large scale PDEs whose scientific output his impacted by this problem. To gain access to dynamics in an exascale simulation that are not identifiable a priori and to make that dynamical data available to machine learning requires fundamental research in the area of in situ data data analytics. Here data analytics includes compression, visualization, uncertainty quantification, and machine learning. This in situ data analytics will enable on-the-fly spatial and temporal compression of solution dynamics, expose that space-time compressed field to machine learning algorithms that have been specialized to work with dynamically evolving data (existing machine learning algorithms treat data sets as static), greatly improving the opportunity for machine learning to provide feedback to the compression, all within an ongoing simulation, without the need to write data to files. The same concepts are also being applied to uncertainty quantification and multi-fidelity modeling which have similar needs for spatial and temporal compression of the ongoing exascale simulation to perform either without the typical, unacceptable writing of data to files.

97 MATHEMATICS AND COMPUTING

ZFP: A compressed array representation for numerical computations

HPC trends favor algorithms and implementations that reduce data motion relative to FLOPS. We investigate the use of lossy compressed data arrays in place of traditional IEEE floating point arrays to store the primary data of calculations. Simulation is fundamentally an exercise in controlled approximation, and error introduced by finite-precision arithmetic (or lossy compression) is just one of several sources of error that need to be managed to ensure sufficient accuracy in a computed result. We describe ZFP, a compressed numerical format designed for in-memory storage of multidimensional arrays, and summarize theoretical results that demonstrate that the error of repeated lossy compression can be bounded and controlled. Furthermore, we establish a relationship between grid resolution and compression-induced errors and show that, contrary to conventional floating point, ZFP reduces finite-difference errors with finer grids. We present example calculations that demonstrate data reduction by 4x or more with negligible impact on solution accuracy. Our results further demonstrate several orders-of-magnitude increase in accuracy using ZFP over IEEE floating point and Posits for the same storage budget.

Lindstrom, Peter

Lifting MGARD: Construction of (pre)wavelets on the interval using polynomial predictors of arbitrary order

MGARD (MultiGrid Adaptive Reduction of Data) is an algorithm for compressing and refactoring scientific data, based on the theory of multigrid methods. The core algorithm is built around stable multilevel decompositions of conforming piecewise linear $C^0$ finite element spaces, enabling accurate error control in various norms and derived quantities of interest. In this work, we extend this construction to arbitrary order Lagrange finite elements $\mathbb{Q}_p$, $p \geq 0$, and propose a reformulation of the algorithm as a lifting scheme with polynomial predictors of arbitrary order. Additionally, a new formulation using a compactly supported wavelet basis is discussed, and an explicit construction of the proposed wavelet transform for uniform dyadic grids is described.

Reshniak, Viktor [Oak Ridge National Laboratory (O

What to Support When You’re Compressing

Over the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research.

Error-Bounded Lossy Compression

StOKeDMD: Streaming Occupation kernel dynamic mode decomposition

Dynamic mode decomposition (DMD) has become a common technique for constructing surrogate models for dynamical systems from observed system states. The Occupation Kernel DMD (OKDMD) method proposed in (Rosenfeld et al., 2022) and (Rosenfeld et al., 2024) is a Liouville operator based method that builds surrogate models from system state trajectories. Here, this paper proposes an extension of OKDMD to the case when the system states are observed in a streaming fashion, i.e., only a small fraction of the state trajectory is available at a given time. The developed method, Streaming Occupation Kernel DMD (StOKeDMD), accommodates the streaming data input by leveraging properties of specific choices of kernel functions and occupation kernels. We apply the StoKeDMD method as a compression method for streaming data, analyze the memory complexity, and demonstrate the performance of StoKeDMD in the compression of streaming data generated from a Lorenz system and a fluid flow simulation.

97 MATHEMATICS AND COMPUTING

Optimising the processing and storage of visibilities using lossy compression

The next-generation radio astronomy instruments are providing a massive increase in sensitivity and coverage, largely through increasing the number of stations in the array and the frequency span sampled. The two primary problems encountered when processing the resultant avalanche of data are the need for abundant storage and the constraints imposed by I/O, as I/O bandwidths drop significantly on cold storage. An example of this is the data deluge expected from the SKA Telescopes of more than 60 PB per day, all to be stored on the buffer filesystem. While compressing the data is an obvious solution, the impacts on the final data products are hard to predict. In this paper, we chose an error-controlled compressor – MGARD – and applied it to simulated SKA-Mid and real pathfinder visibility data, in noise-free and noise-dominated regimes. As the data have an implicit error level in the system temperature, using an error bound in compression provides a natural metric for compression. MGARD ensures the compression incurred errors adhere to the user-prescribed tolerance. To measure the degradation of images reconstructed using the lossy compressed data, we proposed a list of diagnostic measures, exploring the trade-off between these error bounds and the corresponding compression ratios, as well as the impact on science quality derived from the lossy compressed data products through a series of experiments. We studied the global and local impacts on the output images for continuum and spectral line examples. We found relative error bounds of as much as 10%, which provide compression ratios of about 20, have a limited impact on the continuum imaging as the increased noise is less than the image RMS, whereas a 1% error bound (compression ratio of 8) introduces an increase in noise of about an order of magnitude less than the image RMS. For extremely sensitive observations and for very precious data, we would recommend a 0.1% error bound with compression ratios of about 4. These have noise impacts two orders of magnitude less than the image RMS levels. At these levels, the limits are due to instabilities in the deconvolution methods. We compared the results to the alternative compression tool DYSCO, in both the impacts on the images and in the relative flexibility. MGARD provides better compression for similar error bounds and has a host of potentially powerful additional features.

Techniques: interferometric

A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and Visualization

Multi-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to 3.3× under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements.

Wang, Daoce

TensorID v1.0

This Python software package includes new and efficient algorithms for satellite and core interpolative decomposition of tensor data. In general, these algorithms target high-dimensional data reduction and compression. The software is purely numerical and can be applied by others to many important sources of tensor data generated by computation or experiment.

Zhang, Yifan [Lawrence Berkeley National Laborator

Effect of Part Size, Displacement Rate, and Aging on Compressive Properties of Elastomeric Parts of Different Unit Cell Topologies Formed by Vat Photopolymerization Additive Manufacturing

Due to its ability to achieve geometric complexity at high resolution and low length scales, additive manufacturing (AM) has increasingly been used for fabricating cellular structures (e.g., foams and lattices) for a variety of applications. Specifically, elastomeric cellular structures offer tunability of compliance as well as energy absorption and dissipation characteristics. However, there are limited data available on compression properties for printed elastomeric cellular structures of different designs and testing parameters. In this work, the authors evaluate how unit cell topology, part size, the rate of compression, and aging affect the compressive response of polyurethane-based simple cubic, body-centered, and gyroid structures formed by vat photopolymerization AM. Finite element simulations incorporating hyperelastic and viscoelastic models were used to describe the data, and the simulated results compared well with the experimental data. Of the designs tested, only the parts with the body-centered unit cell exhibited differences in stress–strain responses at different part sizes. Of the compression rates tested, the highest displacement rate (1000 mm/min) often caused stiffer compressive behavior, indicating deviation from the quasi-static assumption and approaching the intermediate rate response. The cellular structures did not change in compression properties across five weeks of aging time, which is desirable for cushioning applications. This work advances knowledge on the structure–property relationships of printed elastomeric cellular materials, which will enable more predictable compressive properties that can be traced to specific unit cell designs.

36 MATERIALS SCIENCE

Compressive Response and Energy Absorption of Additively Manufactured Elastomers with Varied Simple Cubic Architectures

Additive manufacturing, and particularly the vat photopolymerization process, enables the fabrication of complex geometries at high resolution and small length scales, making it well-suited for fabricating cellular structures (e.g., foams and lattices). Among these, elastomeric cellular structures are of growing interest due to their tunable compliance and energy dissipation. However, comprehensive data on the compressive behavior of these structures remains limited, especially for investigating the structure-property effects from changing the density and distribution of material within the cellular structure. This study explores how the mechanical response of polyurethane-based simple cubic structures changes when varying volume fraction, unit cell length, and unit cell patterning, which have not been systematically investigated previously in additively manufactured elastomers. Increasing volume fraction from 10% to 50% yielded significant changes in compressive stress–strain performance (decreasing strain at 0.5 MPa by 41.6% and increasing energy absorption density by 3962.5%). Although changing the unit cell length between 2.5 and 7 mm in ~30 mm parts did not result in statistically different stress–strain responses, modifying the configuration of struts of different thicknesses across designs with 30% volume fraction altered the stress–strain behavior (differences of 12.5% in strain at 0.5 MPa and 109.4% for energy absorption density). Power law relationships were developed to understand the interactions between volume fraction, unit cell length, and elastic modulus, and experimental data showed strong fits (R 2 > 0.91). These findings enhance the understanding of how multiple structural design aspects influence the performance of elastomeric cellular materials, providing a foundation for informing strategic design of tailorable materials for diverse mechanical applications.

36 MATERIALS SCIENCE

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science

A surprising proliferation of detwinning in β -tin at extreme loading rates

Integrating data from dynamic compression experiments of condensed matter across three national laboratories has led to insight and quantitative calibration of materials strength over decades of loading rate. For many materials, a single strength model (such as PTW) is sufficient to capture the flow-stress strain rate relationship which is monotonic. Here, we show here that β -tin, a tetragonal metal, exhibits dramatic deviations from this behavior. Naive fitting to a single PTW model is insufficient to capture the behavior; indeed, such resulting inferred flow stress versus strain exhibits a non-monotonic behavior. We suggest a resolution to this by proposing that in β -tin there are important Bauschinger effects arising from favorable conditions for twinning and detwinning. A simple yield surface model when paired with PTW hardening captures the experimental data.

36 MATERIALS SCIENCE

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS