Search NASA⌕ Search

SEARCH · Search NASA

Results for “data compression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

SNAPRed: Reduction of multidimensional neutron time-of-flight diffraction data

SNAP is a neutron time-of-flight diffractometer at the Spallation Neutron Source operated by Oak Ridge National Laboratory. It generates large arrays of neutron detection events that encode the crystalline atomic structure of materials under study. SNAPRed is an application that makes these datasets accessible to end users by orchestrating the process of data reduction while automatically managing the variable neutron instrumentation configuration. It supports arbitrary grouping and masking of individual detector pixels and includes custom-developed data compression approaches to accommodate the large volumes of data generated by the SNAP instrument.

Diffraction↗

buhito

buhito is a Python library for graph analysis and machine learning. Graphs can represent networks with objects as nodes and their relationships as edges. buhito focuses on graphlet methods that study graphs through enumerating their component subgraphs to enable interpretable and fast models of complex systems. The package provides tools for different algorithmic designs for computing, analyzing, and applying graphlets to research problems such as machine learning, data compression, and anomaly detection in graph-structured data. A central feature is performing decomposition data analysis on graphs for machine learning models. Implemented in Python and built upon open-source scientific libraries such as NetworkX, NumPy, and SciPy, buhito provides high-performance methods for researchers exploring the mathematical and computational foundations of graphlet analysis applicable to systems of different sizes.

Pimonova, Yulia↗

Spatial degradation of satellite data

Consideration is given to a technique for spatially degrading high-resolution satellite data to produce comparable data sets over a range of coarser resolutions. Landsat MSS data is used to produce seven spatial resolution data sets by applying a spatial filter designed to simulate sensor response. Also, spatial degradation of coarse resolution data to provide data compression for the production of global-scale data sets is examined. NOAA AVHRR Global Area Coverage data is compared to other sampling procedures. It is found that sampling procedures that incorporate averaging result in decreased variance, while sampling procedures adopting single-value selection have higher variances and produce data values comparable with those from the original data.

Justice, C. O.↗

Method of Real-Time Principal-Component Analysis

Dominant-element-based gradient descent and dynamic initial learning rate (DOGEDYN) is a method of sequential principal-component analysis (PCA) that is well suited for such applications as data compression and extraction of features from sets of data. In comparison with a prior method of gradient-descent-based sequential PCA, this method offers a greater rate of learning convergence. Like the prior method, DOGEDYN can be implemented in software. However, the main advantage of DOGEDYN over the prior method lies in the facts that it requires less computation and can be implemented in simpler hardware. It should be possible to implement DOGEDYN in compact, low-power, very-large-scale integrated (VLSI) circuitry that could process data in real time.

Duong, Tuan↗

Real-Time Principal-Component Analysis

A recently written computer program implements dominant-element-based gradient descent and dynamic initial learning rate (DOGEDYN), which was described in Method of Real-Time Principal-Component Analysis (NPO-40034) NASA Tech Briefs, Vol. 29, No. 1 (January 2005), page 59. To recapitulate: DOGEDYN is a method of sequential principal-component analysis (PCA) suitable for such applications as data compression and extraction of features from sets of data. In DOGEDYN, input data are represented as a sequence of vectors acquired at sampling times. The learning algorithm in DOGEDYN involves sequential extraction of principal vectors by means of a gradient descent in which only the dominant element is used at each iteration. Each iteration includes updating of elements of a weight matrix by amounts proportional to a dynamic initial learning rate chosen to increase the rate of convergence by compensating for the energy lost through the previous extraction of principal components. In comparison with a prior method of gradient-descent-based sequential PCA, DOGEDYN involves less computation and offers a greater rate of learning convergence. The sequential DOGEDYN computations require less memory than would parallel computations for the same purpose. The DOGEDYN software can be executed on a personal computer.

Duong, Vu↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Joint pattern recognition/data compression concept for ERTS multispectral imaging

This paper describes a new technique which jointly applies clustering and source encoding concepts to obtain data compression. The cluster compression technique basically uses clustering to extract features from the measurement data set which are used to describe characteristics of the entire data set. In addition, the features may be used to approximate each individual measurement vector by forming a sequence of scalar numbers which define each measurement vector in terms of the cluster features. This sequence, called the feature map, is then efficiently represented by using source encoding concepts. A description of a practical cluster compression algorithm is given and experimental results are presented to show trade-offs and characteristics of various implementations. Examples are provided which demonstrate the application of cluster compression to multispectral image data of the Earth Resources Technology Satellite.

Hilbert, E. E.↗

Generation and Performance of Automated Jarosite Mineral Detectors for Vis/NIR Spectrometers at Mars

Sulfate salt discoveries at the Eagle and Endurance craters in Meridiani Planum by the Mars Exploration Rover Opportunity have proven mineralogically the existence and involvement of water in Mars past. Visible and near infrared spectrometers like the Mars Express OMEGA, the Mars Reconnaissance Orbiter CRISM and the 2009 Mars Science Laboratory Rover cameras are powerful tools for the identification of water-bearing salts and other high priority minerals at Mars. The increasing spectral resolution and rover mission lifetimes represented by these missions currently necessitate data compression in order to ease downlink restrictions. On board data processing techniques can be used to guide the selection, measurement and return of scientifically important data from relevant targets, thus easing bandwidth stress and increasing scientific return. We have developed an automated support vector machine (SVM) detector operating in the visible/near-infrared (VisNIR, 300-2500 nm) spectral range trained to recognize the mineral jarosite (typically KFe3(SO4)2(OH)6), positively identified by the Mossbauer spectrometer at Meridiani Planum. Additional information is included in the original extended abstract.

Gilmore, M. S.↗

Nisar L-band Digital Electronics Subsystem

The NASA-ISRO Synthetic Aperture Radar (NISAR) L-band SAR instrument employs multiple digital channels to optimize resolution while keeping a large swath on a single pass. High-speed digitization with fine synchronization and digital beam forming are necessary in order to facilitate this new technique called SweepSAR. An architecture employing multiple FPGA based digital signal processors has been conceived to facilitate digital calibration on an individual channel basis as well as digital signal processing to optimize the receive signal. On-board processing and data compression has been implemented to reduce the volume of data in order to satisfy the operational requirements of near global coverage for the desired science targets. A novel command and timing architecture was developed to manage this complex system to meet the challenging project requirements. The NISAR L-band Digital Electronics Subsystem is the combination of the hardware, firmware and software components architected and implemented to operate this radar and return the desired quantity and quality of data for the science community.

SweepSAR↗

NISAR L-SAR Digital Electronics Subsystem - A Multichannel Distributed Processing System with Synchronous Timing Control for Digital Beam Forming and Multiple Echo Tracking

The NASA-ISRO Synthetic Aperture Radar (NISAR) L-band SAR instrument employs multiple digital channels to optimize resolution while keeping a large swath on a single pass. High-speed digitization with fine synchronization and digital beam forming are necessary in order to facilitate this new technique called SweepSAR. An architecture employing multiple FPGA based digital signal processors has been conceived to facilitate digital calibration on an individual channel basis as well as digital signal processing to optimize the receive signal. On-board processing and data compression has been implemented to reduce the volume of data in order to satisfy the operational requirements of near global coverage for the desired science targets. A novel command and timing architecture was developed to manage this complex system while providing detailed control of individual channel receive window timing required for digital beam forming. The NISAR L-band Digital Electronics Subsystem is the combination of the hardware, firmware and software components architected and implemented to operate this radar and return the desired quantity and quality of data for the science community.

Chuang, Chung-Lun↗

Adaptive coding of MSS imagery

A number of adaptive data compression techniques are considered for reducing the bandwidth of multispectral data. They include adaptive transform coding, adaptive DPCM, adaptive cluster coding, and a hybrid method. The techniques are simulated and their performance in compressing the bandwidth of Landsat multispectral images is evaluated and compared using signal-to-noise ratio and classification consistency as fidelity criteria.

Habibi, A.↗

Compressive Sensing Based Data Acquisition Architecture for Transient Stellar Events in Crowded Star Fields

Compressive sensing is a mathematical technique for simultaneous data acquisition and compression. In this work, we show a CS based architecture for acquiring and reconstructing transient stellar events. This architecture recon-structs a differenced image itself, eliminating the need for any sparse domain transforms, otherwise required for traditional CS reconstruction. The resulting reconstructed differenced image is of critical importance as the information required for generating a time-series photometric light curve is obtained only from a differenced image. Hence, reconstructing a crowded star spatial image, followed by differencing is wasteful. This architecture eliminates the need to 1.) transform an image to a sparse domain, 2.) Reconstruct a dense field, and then apply differencing on the image to obtain the image of critical value. We study the case of microlensing to depict a star source experiencing magnification in time. Our results show that this architecture is able to reconstruct star source magnitudes with magnification factors greater than 1 for clean images with error less than 2% using only 10% of the Nyquist rate samples.

Compressive Sensing, Data acquisition, Image diffe↗

Study of efficient video compression algorithms for space shuttle applications

Results are presented of a study on video data compression techniques applicable to space flight communication. This study is directed towards monochrome (black and white) picture communication with special emphasis on feasibility of hardware implementation. The primary factors for such a communication system in space flight application are: picture quality, system reliability, power comsumption, and hardware weight. In terms of hardware implementation, these are directly related to hardware complexity, effectiveness of the hardware algorithm, immunity of the source code to channel noise, and data transmission rate (or transmission bandwidth). A system is recommended, and its hardware requirement summarized. Simulations of the study were performed on the improved LIM video controller which is computer-controlled by the META-4 CPU.

Poo, Z.↗

A Framework for Compressing Unstructured Scientific Data via Serialization

We present a general framework for compressing unstructured scientific data with known local connectivity. A common application is simulation data defined on arbitrary finite element meshes. The framework employs a greedy topology preserving reordering of original nodes which allows for seamless integration into existing data processing pipelines. This reordering process depends solely on mesh connectivity and can be performed offline for optimal efficiency. However, the algorithm’s greedy nature also supports on-the-fly implementation. The proposed method is compatible with any compression algorithm that leverages spatial correlations within the data. The effectiveness of this approach is demonstrated on a large-scale real dataset using several compression methods, including MGARD, SZ, and ZFP.

Reshniak, Viktor [ORNL] (ORCID:0000000315454462)↗