Search NASA⌕ Search

SEARCH · Search NASA

Results for “DATA COMPRESSION”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Cosmological constraints from a joint DESI DR1 Full-Shape and DR2 BAO

We present a cosmological analysis combining full-shape (FS) clustering measurements from the Dark Energy Spectroscopic Instrument (DESI) DR1 with baryon acoustic oscillation (BAO) measurements from DESI DR2. To achieve a robust combination that accounts for the correlation between the two data releases, we employ the ShapeFit compression method and estimate the joint covariance using EZmocks. This compressed approach inherently mitigates the prior volume effects that have previously dominated Bayesian constraints from DESI data with minimal external priors. Consequently, we obtain — for the first time within a Bayesian framework — reliable DESI-only constraints on extensions to ΛCDM using only a Big Bang Nucleosynthesis prior on the baryon density and a wide prior on the spectral index. In flat ΛCDM, we find Ω m = 0.3035 ± 0.0085, h = 0.6876 ± 0.0059, and σ 8 = 0.822 ± 0.034. For the w 0 w a CDM dynamical dark energy model, we measure w 0 = -0.49 ± 0.25 and w a = -1.52 ± 0.77, improving constraints by ∼ 30% relative to the analogous DR1 measurement and reducing the discrepancy with ΛCDM to 1.4σ when compared to BAO only analyses. We also report competitive limits on the sum of neutrino masses and spatial curvature. This work demonstrates that the ShapeFit compression provides a prior-robust and computationally efficient pathway to constrain beyond-ΛCDM physics with large-scale structure.

baryon acoustic oscillations↗

HP-MDR: High-performance and Portable Data Refactoring and Progressive Retrieval with Advanced GPUs

Scientific applications produce vast amounts of data, posing grand challenges in the underlying data management and analytic tasks. Progressive compression is a promising way to address this problem, as it allows for on-demand data retrieval with significantly reduced data movement cost. However, most existing progressive methods are designed for CPUs, leaving a gap for them to unleash the power of today’s heterogeneous computing systems with GPUs.In this work, we propose HP-MDR, a high-performance and portable data refactoring and progressive retrieval framework for GPUs. Our contributions are four-fold: (1) We carefully optimize the bitplane encoding and lossless encoding, two key stages in progressive methods, to achieve high performance on GPUs; (2) We propose pipeline optimization and incorporate it with data refactoring and progressive retrieval workflows to further enhance the performance for large data process; (3) We leverage our framework to enable high-performance data retrieval with guaranteed error control for common Quantities of Interest; (4) We evaluate HP-MDR and compare it with state of the arts using five real-world datasets. Experimental results demonstrate that HP-MDR delivers an average 13.68 × and 6.31 × throughput in data refactoring and progressive retrieval tasks, respectively. It also leads to 11.22 × throughput for recomposing required data representations under Quantity-of-Interest error control and 6.04 × performance for the corresponding end-to-end data retrieval, when compared with state-of-the-art solutions.

Li, Yanliang [University of Oregon]↗

Shock compression of liquid helium to 360 GPa

Data for the shock equation of state of helium are obtained up to 360⁢G⁡Pa, nearly doubling the pressure of previous experimental measurements. The helium samples are first precompressed to 2.7 g⁡c⁢m −3 in a diamond anvil cell prior to laser-driven shock compression at the Omega Laser Facility. Time-resolved Doppler velocimetry and pyrometry reveal significant reflectivity and greater compressibility compared to the predictions of existing broad-range tabular equation of state models, which may be caused by the onset of ionization. These experimental observations, however, are largely captured with molecular-dynamics simulations based on density functional theory, affirming the ability of first-principles techniques to capture complex physics, while enabling critical insight for the behavior of warm dense helium in Jovian interiors and white dwarf atmospheres.

Physics - Plasma physics↗

Initial Testing of an In Situ Load Retention Aging Vessel

A thermal aging vessel instrumented with load cells was fabricated. The primary function of the vessel is to continuously monitor the in situ load retention of up to three compressed polymer coupons undergoing thermally accelerated aging under nitrogen. A secondary function is to enable gas sampling of the vessel headspace during thermal aging. Heating of the vessel is achieved using a custom heater jacket. To improve upon our conventional aging study methods which require periodic interruption of aging to perform load testing in an Instron machine at room temperature, this technology aims to automate/facilitate data acquisition/analysis, improve data quality, and enable uninterrupted compression of the polymer which represents the service condition. As an example case to assess functionality of the in situ vessel, the load retention of a siloxane elastomer material additively manufactured by direct-ink-writing (DIW) was measured at three different isothermal aging temperatures for ~1 month. Initial compression of the coupons while near the aging temperature was achieved by temporarily opening the heated vessel to access the interior chamber and manually tightening four nuts to drive the heated compression plate down onto the heated coupons. Initial testing demonstrated achievement of the primary load retention monitoring function. Unfortunately, the vessel leaked which prevented gas sampling; an active purge was used to maintain a nitrogen atmosphere. Welded or otherwise sealed joints, which could be implemented in a future design, would likely eliminate leak paths. To apply time-temperature superposition (TTS), a technique used to provide long-term prediction of the load retention from short-term isothermal data, the load retention needed to be calculated relative to the load at an estimated “equilibrium” time, after most of the transient viscoelastic physical relaxation occurred. The peak load immediately after compression could not be used as the load retention basis for two reasons: (1) age-related changes must be isolated from non-age-related physical relaxation before applying TTS and (2) the manual mechanism used to compress the specimens at the aging temperature was neither smooth nor repeatable which affected the peak load value. To better understand the effect of the mode of initial compression on the measured load, and possibly better estimate “equilibrium” physical relaxation times, systematic stress relaxation experiments were performed using an Instron machine with a thermal chamber. At a given temperature, the DIW polymer was compressed to a fixed strain in either a stepped or continuous manner at two different rates, then held at that strain for 24 hrs. The results indicated that, at a given temperature, the different stress relaxation curves appeared to converge to the same curve at some “equilibrium” time when the non-age-related physical relaxation was mostly complete. Though this observation suggests that the discontinuous manual compression employed by the vessel is feasible, a compression mechanism that is rapid, smooth, and repeatable would enhance its use.

36 MATERIALS SCIENCE↗

Synthetic Streamflow Datasets to Support Emulation of Water Allocations via LSTM

This archive is the data companion to the bonney_et-al_2026_erc metarepo which generates synthetic data, trains an LSTM model, and generates performance metrics on the trained model. While the generation of the synthetic data is fully reprodicible, it is a computationally expensive process. This data archive contains the synthetic datasets needed for training and testing an LSTM model and reproduction of figures and tables. In addition, supplemenatary data products generating and visualizing results is also included, such as geospatial data for the basin. Contents There are two high level directories: `WRAP_archive/` and `repo_data/`. The `WRAP_archive` directory contains compressed intermediate dataproducts from the dataset generation workflow (marked as "I_Dataset_Generation" in the metarepo). These data products are not required by any scripts in the metarepo, but they are archived as they are expensive to generate and may have useful information for other analyses. The `repo_data` directory contains the necessary data for reproducing the workflow in the metarepo and should be decompressed and moved into the top level of the metarepo. Additional details are provided in README.md.

drought↗

Leafweb: Leaf Gas Exchange and Pulse-Amplitude Modulated Fluorometry for C4 Species, June 2026 Release

This dataset contains leaf gas exchange and Pulse-Amplitude Modulated (PAM) fluorometry for 98 C4 species. The C4 photosynthetic pathway employs specialized CO2 concentration mechanisms and Kranz anatomy to enrich CO2 concentration around Rubisco, the enzyme that catalyzes carbon fixation in the Calvin-Benson cycle to suppress photorespiration and increase the use efficiencies of light, nitrogen, and water as compared to the C3 photosynthetic pathways. Large-scale C4 photosynthetic datasets are relatively scarce, which has affected C4 photosynthesis research. To improve C4 photosynthetic data availability, Leafweb organized an effort to systematically collect, compile, standardize, and organize measurements of leaf gas exchange and/or Pulse-Amplitude Modulated (PAM) fluorometry of C4 species. This derived a C4 photosynthetic dataset containing measurements made by independent researchers in multiple countries in various environments (field, garden, or greenhouse). It covers three biochemical subtypes – the nicotinamide adenine dinucleotide phosphate-malic enzyme (NADP-ME), nicotinamide adenine dinucleotide-malic enzyme (NAD-ME), and phosphoenolpyruvate carboxykinase (PEP-CK) subtypes. This dataset is useful for using Artificial Intelligence / Machine Learning and mechanistic models to study C4 photosynthesis and compare across different biochemical subtypes. This dataset contains 3 compressed (*.zip) folders containing 1,892 data files in comma-separate values (*.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma-separate values (*.csv) format and a user guide in PDF (*.pdf) format.

Zhou, Haoran [Tianjin University, China]↗

NLR HPC Facility Power Usage Effectiveness (PUE) Data

Timeseries of Energy Systems Integration Facility (ESIF) Data Center Power Usage Effectiveness (PUE) Data provided in Parquet and compressed CSV formats Power Metrics Timeseries Fields: ts: Timestamp cooling_kw: Cooling (kilowatts) - Captures the power used by fans and pipe trace heaters associated with outdoor cooling equipment. The dedicated tower filter pump power is also captured as cooling load. energy_reuse: Energy Reuse Effectiveness hvac_kw: Heating, ventilation, and air conditioning (kilowatts) - Captures fan walls, fan coils that support the data center electrical rooms, and the make-up air unit. it_power_kw: IT equipment (kilowatts) - Captures power used by the IT equipment on the data center floor. plug_and_light_kw: Lights and utility plugs (kilowatts) - Captures power associated with the data center and dedicated mechanical room. The crank-case heater for the emergency standby generator is also captured as light and plug load. pue: Power Usage Effectiveness pump_kw: Pumps (kilowatts) - Captures power from pumps that move water in the data center Energy Recover Water loop and the Tower Water loops, and also captures power used by the boost pumps that circulate water through the fan walls. Note: The tower filter pump runs constantly to filter water from the data center cooling tower system, so 2.67 kilowatts are attributed to this pump and that is not reflected in this data field. day: Day of month Outside Weather Station Timeseries Fields: ts: Timestamp outside_air_humidity: Outside air humidity - Relative humidity percent outside_air_temp: Outside air temperature - Degrees Fahrenheit day: Day of month More detail: High-Performance Computing Data Center Power Usage Effectiveness

97 MATHEMATICS AND COMPUTING↗

Evaluation of Drilling Performance at The Geysers with Machine Learning Methods Using Geologic Data

A recent well, GDC-36, was drilled in The Geysers Geothermal Field served in a Department of Energy-industry to demonstrate improved drilling performance with polycrystalline diamond compact (PDC) bits. Both PDC and roller cone drill bits were used to drill this well. Key challenges encountered during drilling included lost circulation in the mud-drilled section, and bit damage interfacial severity in the deeper, air-drilled section. The objective of this study is to evaluate the drilling performance in relation to the local geological characteristics using machine learning methods. By applying K-clustering to the sonic log data, we were able to identify areas correlated with measured lost circulation. Also, the boundaries defined by clustering of the mineralogical and lithological data from the mud logs correlate well with interfacial severity during drilling. A random forest model was employed to build correlation between drilling data and rock strength. The confined compressive strength (CCS) of the rock in the training of the machine learning model was inferred from the dipole sonic log. The R-squared of the testing data is 0.78, and the RMSE (Root Mean Squared Error) is 0.06. The trained model was used to forecast rock strength for the section where sonic log data are not available. CCS could also be inferred from mud logs provided the relationship between mineralogy and rock strength is established through core testing data.

15 GEOTHERMAL ENERGY↗

Compression–tension cell with sample manipulator for in situ X‐ray nanotomography experiments

In situ X-ray nanotomography experiments where tensile or compressive force is applied on the sample require specialized equipment. A compression-tension device with fluid flow-through capability has been designed for X-ray nanotomography beamlines. The compression-tension cell is equipped with a triaxial stage for sample alignment and a high sensitivity loadcell for measurement of applied force. To handle the <100 µm samples used for X-ray nanotomography imaging and for loading samples on the compression-tension cell a sample manipulator has been built. The sample manipulator is capable of selecting a single <100 µm particle for nanotomography scanning while viewing multiple samples under an optical microscope. To test the functionality of these two devices an initial compression experiment involving two glass beads was performed. To demonstrate instrument stability two spherical glass beads were compressed from a no load condition until one of the beads fractured. Nanotomography data were collected at each step of increasing compressive force. The experimentally observed contact area of the spherical glass beads was compared with the theoretical estimate using the Hertz analysis. To demonstrate the fluid flow capability, two calcite grains were compressed against each other under a calcite saturated solution. Surface topological changes were observed for the stressed grain contact area.

X-ray tomography↗

Direct observation of diamond formation in a shock-compressed high explosive

Understanding the formation timescale and structure of carbonaceous reaction products is critical for modeling the high-pressure equation-of-state of organic materials. We use the National Ignition Facility to shock-compress polycrystalline TATB (C 6 H 6 N 6 O 6 ) samples to ~70–130 GPa and ~4000–5500 K, employing in situ nanosecond X-ray diffraction to probe reaction products and velocimetry to measure transmitted compression wave profiles. Our diffraction data is consistent with the formation of diamond over timescales less than ~60 ns. This represents carbon condensation from a molecular explosive on timescales three times faster than previously reported and the earliest observation of diamond produced from reacting TATB. Reactive flow simulations with explicit chemistry reproduce the observed temporal structure within wave profiles to inform the distribution of P-T states. These findings provide direct evidence of ultrafast diamond formation in a reactive system at extreme conditions and provide new constraints for models of shock and detonation chemistry.

Clarke, Samantha M. [Lawrence Livermore National L↗

Integration of RNTuple in ATLAS Athena

After using ROOT’s TTree I/O subsystem for over two decades and storing more than an exabyte of compressed High Energy Physics (HEP) data, advances in technology have motivated a complete redesign, RNTuple, which breaks backward-compatibility to take better advantage of these storage options. The RNTuple I/O subsystem has been designed to address performance bottlenecks and other shortcomings of TTree. Specifically, RNTuple comes with an updated, more compact binary data format that can be stored both in ROOT files and natively in object stores. It is designed for modern storage hardware (e.g. high-throughput low-latency NVMe SSDs), and provides robust and easy to use interfaces. The binary format of RNTuple is scheduled to become production grade in 2024, and recently has become mature enough to start exploring the integration into software used by HEP experiments. In this contribution, we discuss the developments to support the features as required by the ATLAS analysis Event Data Model (EDM) in RNTuple, which will enable its integration into the Athena software framework. With these developments in place, we evaluate the performance of the current most recent versions of RNTuple-based ATLAS data sets and compare this to that of TTree.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

Elucidating texture and grain morphology contributions to the micromechanical response of additively manufactured Inconel 625

Microstructural variation of additively manufactured (AM) metal components in comparison to wrought counterparts makes certification for critical applications a challenge. Microscale simulations leveraging modern computational tools may be used to supplement testing of AM microstructures, thus accelerating certification by reducing the number of experiments needed. However, as micromechanical response is closely tied to critical properties like fatigue-life and fracture, utilization of these simulations with macroscale experimental data alone is insufficient. One means to attain microscale experimental data is in situ diffraction data collected from synchrotron X-ray sources. In this work, such data were collected during in situ compression of AM Inconel 625 superalloy. Interpretation of experimental results was assisted by massive (8M element) complementary micromechanical simulations performed on sets of virtual microstructures generated using cellular automata. Together, micromechanical data from diffraction experiments and simulations were used to probe the effects of textured “track” microstructures generated during laser powder bed fusion and directional strength-to-stiffness on micromechanical response. Though fiber-averaged directional strength-to-stiffness ratios were expected to dominate given the high elastic anisotropy of the material, the combination of small variations in texture and specific grain configurations unique to AM microstructures lead to significant variability in micromechanical response after yield. The findings emphasize the importance of high-fidelity microstructural representation that captures key texture components and AM-specific morphology for property prediction of AM metals.

36 MATERIALS SCIENCE↗

Refractive index of lithium fluoride at high dynamic stresses

Alkali halides are prototypical ionic solids whose refractive index is a fundamental property related to their lattice structure and ionic polarizability. In particular, lithium fluoride (LiF) has the largest band gap of any known transparent material and maintains transparency at high pressures, making it well suited for use as an optical window in dynamic compression experiments. While empirical fits to the measured density dependence of the refractive index have been provided, a model that is valid over the entire pressure range of experiments and based on a polarization-based theoretical description is lacking. We present a refractive index model for dynamically compressed LiF based on the Lorentz-Lorenz equation, where the molecular polarizability is determined using a single oscillator model with a strain polarizability parameter Λ=0.73. Here, we show that our modeling approach provides an excellent match to the LiF refractive index data for both shock and ramp compression experiments to 900 GPa (density compression greater than threefold). Additionally, updated fits are provided to determine the refractive index correction for LiF windows used in dynamic compression experiments.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

3D Gaussian Splatting for Volume Compression

This codebase uses machine learning to train a collection of 3D Gaussian distributions to approximate scientific volume data. Because this collection uses less memory than the original dataset, it can be used as a compressed model of the original data for applications such as visualization.

Dyken, Landon↗

SPRUCE Whole Ecosystem Warming (WEW) Environmental Data and Water Table Summaries, Marcell Experimental Forest, Minnesota, 2015-2024

This data set contains observations of photosynthetically active radiation (PAR), precipitation, soil temperature, soil volumetric water content, air temperature, relative humidity, and normalized water table depth that are summarized on a daily, weekly, monthly, and annual basis for each of the SPRUCE plots. Observations span 2015-2024. This dataset draws on several datasets (Hanson et al. 2016; Hanson et al. 2020; and Warren, unpublished data) and compiles these environmental observations into useful formats for data analysis. These environmental metrics can be used to understand the environmental conditions inside SPRUCE environmental chambers throughout the durations of the experiment and can be paired with other data for modeling and analysis. R code used to generate these files is provided as part of the data package. This dataset contains four data files in comma separate (.csv) format and a compressed folder (*.zip) containing three R (*.r) scripts. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. User note: Users must cite the original dataset/s along with this dataset when publishing any analyses using this dataset. Details on the dataset used to compile each variable are available in the header row of the files and in the user guide.

air temperature↗

High‐speed 4‐dimensional scanning transmission electron microscopy using compressive sensing techniques

Abstract Here we show that compressive sensing allows 4‐dimensional (4‐D) STEM data to be obtained and accurately reconstructed with both high‐speed and reduced electron fluence. The methodology needed to achieve these results compared to conventional 4‐D approaches requires only that a random subset of probe locations is acquired from the typical regular scanning grid, which immediately generates both higher speed and the lower fluence experimentally. We also consider downsampling of the detector, showing that oversampling is inherent within convergent beam electron diffraction (CBED) patterns and that detector downsampling does not reduce precision but allows faster experimental data acquisition. Analysis of an experimental atomic resolution yttrium silicide dataset shows that it is possible to recover over 25 dB peak signal‐to‐noise ratio in the recovered phase using 0.3% of the total data. Lay abstract : Four‐dimensional scanning transmission electron microscopy (4‐D STEM) is a powerful technique for characterizing complex nanoscale structures. In this method, a convergent beam electron diffraction pattern (CBED) is acquired at each probe location during the scan of the sample. This means that a 2‐dimensional signal is acquired at each 2‐D probe location, equating to a 4‐D dataset. Despite the recent development of fast direct electron detectors, some capable of 100kHz frame rates, the limiting factor for 4‐D STEM is acquisition times in the majority of cases, where cameras will typically operate on the order of 2kHz. This means that a raster scan containing 256^2 probe locations can take on the order of 30s, approximately 100‐1000 times longer than a conventional STEM imaging technique using monolithic radial detectors. As a result, 4‐D STEM acquisitions can be subject to adverse effects such as drift, beam damage, and sample contamination. Recent advances in computational imaging techniques for STEM have allowed for faster acquisition speeds by way of acquiring only a random subset of probe locations from the field of view. By doing this, the acquisition time is significantly reduced, in some cases by a factor of 10‐100 times. The acquired data is then processed to fill‐in or inpaint the missing data, taking advantage of the inherently low‐complex signals which can be linearly combined to recover the information. In this work, similar methods are demonstrated for the acquisition of 4‐D STEM data, where only a random subset of CBED patterns are acquired over the raster scan. We simulate the compressive sensing acquisition method for 4‐D STEM and present our findings for a variety of analysis techniques such as ptychography and differential phase contrast. Our results show that acquisition times can be significantly reduced on the order of 100‐300 times, therefore improving existing frame rates, as well as further reducing the electron fluence beyond just using a faster camera.

Robinson, Alex W.↗