What to Support when You're Compressing: The State of Practice, Gaps, and Opportunities for Scientific Data Compression
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
A suite of experimental infrastructure projects has been developed by the National Reactor Innovation Center to accelerate advanced reactor demonstrations and facilitate their development, addressing crucial gaps in data, materials characterization, and modeling. First, the Molten Salt Thermophysical Examination Capability (MSTEC) provides a specialized platform for post-irradiation characterization of molten salt reactor fuel, coolant salts, and structural materials, essential for supporting the design and operation of advanced reactors and future commercial molten salt reactor development and licensing. The Virtual Test Bed (VTB) complements these efforts by leveraging advanced modeling and simulation tools to evaluate reactor performance and safety. Serving as a library of reference models, the VTB offers a database of multiphysics reactor models, facilitating rapid safety evaluations and includes continuous software quality assurance, crucial for accelerating deployment while maintaining reliability. Additionally, the Helium Component Test Facility (HeCTF) addresses the need for high-temperature helium-cooled reactor component testing. As the first-of-its-kind facility in the United States, HeCTF emulates high-temperature gas reactor conditions, reducing time and cost associated with component validation, thereby accelerating reactor development. Finally, In-cell Thermal Creep Frames provide a unique solution for obtaining thermal creep data from irradiated materials, critical for materials qualification and licensing. Developed by the National Reactor Innovation Center, these compact frames enable the examination of previously irradiated materials, overcoming traditional limitations and enhancing the understanding of mechanical properties crucial for reactor development. Collectively, these experimental infrastructure projects form a comprehensive framework aimed at expediting advanced reactor demonstrations, fostering innovation, and ensuring the viability of next-generation nuclear energy solutions.
The gas and tar composition of a diffusion flame from longleaf pine needles is currently poorly understood and more data are needed to fill in the gap between pyrolysis data and smoke plume data, thus improving physical and chemical modeling of wildland smoke formation. A pilot experiment to measure light gas and tar composition of such a flame is described for three flame regions: persistent flame (flame base), intermittent flame, and smoke plume. Flame gases from 24 experimental fires were collected in canisters and analyzed using EPA method TO-14A for CO 2 , CO, H 2 , CH 4 , and C 2 to C 7 hydrocarbon gases. Condensed gas (tar) samples were collected and analyzed using GC/MS. Other light gases were measured using FTIR spectroscopy. Results from compositional data analysis suggest significant differences in (relative) concentration of compounds detected in the three regions of the flame. Statistical tests for differences in flame zones were performed using the canister data: Concentration of hydrocarbons relative to CO and CO 2 decreased from the persistent flame zone above the pyrolyzing needles through the intermittent flame region into the flame-free plume. This was likely due to both chemical reactions (oxidation) occurring in the flame as well as the introduction of air into the flame/plume by entrainment.
Application energy optimization in HPC data centers face two critical gaps. Systematic methodologies that connect data center policies to application decisions and accessible monitoring tools that enable data-driven optimization. We address both gaps through two complementary pillars. First, we present a methodology based on extended weighted Energy Delay Product (EDP) to translate data center operational priorities and integrate energy considerations into the energy optimization workflow which starts from continuous monitoring through targeted optimization. Second, we present a user-space monitoring tool, Omnistat, that enables this methodology by providing developers with direct access to actionable energy telemetry. Through deployment on the Frontier supercomputer and case studies exploring performance-energy trade-offs, we show how these pillars help energy as an integral optimization target for developers as active participants in data center efficiency.
One of the missions of the US Department of Energy’s Office of Nuclear Energy (DOE-NE) Molten Salt Reactor (MSR) Campaign under the Advanced Reactor Technology program has been to experimentally measure thermophysical properties of MSR-relevent salt systems, with the intent of supporting the development of the Molten Salt Thermal Properties Database (MSTDB). This database is jointly funded by the DOE-NE Nuclear Energy Advanced Modeling and Simulation Program and the MSR Campaign. Multiple DOE national laboratories, including Oak Ridge National Laboratory (ORNL), have been conducting measurements of thermophysical properties to support MSTDB development and provide MSR developers with access to new data that has been measured using modern methodologies and more advanced sample characterization techniques. These data may either fill gaps in the database or provide updated higher quality data to replace legacy data. Researchers at ORNL have recognized significant gaps in the transport property data of actinide-bearing fluoride salt systems of MSR industry interest. Moreover, for the data present in MSTDB in this category, the uncertainty margins are generally high, leading to questionability in our current understanding of the thermophysical characterization of actinide fluoride mixtures. As such, the focus of this study has been to generate new transport property data of actinide fluoride mixtures that are of immediate interest to MSR developers. Specifically, the mixtures NaF-UF 4 (78 - 22 mol%) and NaF-KF-UF 4 (57-16.04-26.91 mol%) have been studied—NaF-UF 4 for thermal conductivity and viscosity and NaF-KF-UF 4 for viscosity. Thermal conductivity measurements have been conducted with a variable gap apparatus, whereas viscosity has been measured with a rolling ball viscometer. Methodological and calibration details are provided for both measurement processes, along with measurement system updates that have enabled easier manufacturing of components and fewer challenges associated with conducting the measurements themselves. The resultant data collected for NaF-UF 4 (78–22 mol%) and NaF-KF-UF 4 (57-16.04-26.91 mol%) have been compared with relevant mixture data within the thermophysical arm of the MSTDB (MSTDB-TP).
This dataset contains 15-minute tide height and salinity data from the Typha site along the Parker River, part of the Plum Island Ecosystems Long Term Ecological Research (PIE LTER) site in Plum Island Sound, Massachusetts (MA) 2014-2023. Tide height (in NAVD88) was compiled from measurements conducted at the mouth of Plum Island Sound and corrected for time lags. Gap-filling of missing periods were done by fitting tidal constituents to the time series. Salinity was measured (and is stored on ESS DIVE ) in 2022 and 2023 using HOBO U24-002 conductivity loggers. River discharge is the most important control on tidal river water salinity at the location (Vallino & Hopkinson, 1998). An artificial neural network was trained to predict river water salinity at the location using Parker River discharge (USGS station 01101000, Parker River at Byfield, MA) and gap-filled salinity observations from a long-term monitoring station ca. 3km downstream from the Typha site (LTER station ‘Middle Road’) as input variables to create continuous time series information. The data set was used in the spin up and simulations of a land surface model coupled to a biogeochemical reaction network (ELM PFLOTRAN) assessing impacts of hydrology and salinity input on methane fluxes in 2022 and 2023 (Sulman et al., 2024). Metadata files ELMPFLOTRAN_tide_salinity_dd.csv and ELMPFLOTRAN_tide_salinity_flmd.csv provide details on site location, data variables, and QA/QC methods .
This data contains output from the pumphouse eddy covariance tower that includes shortwave radiation, longwave radiation, net radiation, air temperature, relative humidity, as well as sensible, latent, and ground heat fluxes. Also included is calculated evapotranspiration from the latent heat flux and the latent heat of vaporization. All data are on a daily timestep and displayed in Mountain Time. The data has been processed, and Quality Assurance / Quality Control (QA/QC) was done, but any daily gaps in the data have not been filled in. This research was funded by the Department of Energy and performed as part of the Watershed Function Scientific Focus Area. This research aimed to constrain evapotranspiration in a high-elevation catchment.The dataset includes one comma-separated values (CSV) data file (EddyCovariance_MeteorlogicalVariables_CrestedButtePumphouse.csv). Additionally, three metadata CSV files are included: (1) location metadata file (locations.csv), which contains location metadata and coordinates; (2) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and (3) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.
Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.
As Automated Vehicle (AV) services proliferate, data sharing between AV operators and municipal agents is assuming greater importance. Information on the dynamic nature of the road system such as incidents to avoid, weather hazards (such as flooding), construction and detours, as well as active safety concerns (e.g. - riots) is important for AV operators. Such information cannot be directly sensed from a vehicle's sensor array, but instead must be communicated in a timely and trustworthy channel. Municipalities are interested in pushing this information to AV operators to support emergency response efforts, reduce traffic in construction zones, and generally improve operation of the system. Similarly, information on vehicle safety such as disengagements, as well as critical information on the use of roadway system (trips, origin and destination patterns) are important performance factors for municipalities to understand utilization and plan for appropriate infrastructure. As mobility shifts to on-demand options, the need for safe and coordinated pick-up and drop-off zones will increase (potentially reducing parking needs). For all of these reasons, communication flows between AV operators and municipalities are becoming increasingly important. This paper investigates the functions, emerging practices and protocols for sharing of such critical data, and identifies gaps in and challenges in existing practices. Additionally, case studies are used to highlight the impacts of data sharing between AV operators and municipalities.
Abstract Given the pressing challenges posed by climate change, it is crucial to develop a deeper understanding of the impacts of escalating drought and heat stress on terrestrial ecosystems and the vital services they offer. Soil and plant water potential play a pivotal role in governing the dynamics of water within ecosystems and exert direct control over plant function and mortality risk during periods of ecological stress. However, existing observations of water potential suffer from significant limitations, including their sporadic and discontinuous nature, inconsistent representation of relevant spatio-temporal scales and numerous methodological challenges. These limitations hinder the comprehensive and synthetic research needed to enhance our conceptual understanding and predictive models of plant function and survival under limited moisture availability. In this article, we present PSInet (PSI—for the Greek letter Ψ used to denote water potential), a novel collaborative network of researchers and data, designed to bridge the current critical information gap in water potential data. The primary objectives of PSInet are as follows. (i) Establishing the first openly accessible global database for time series of plant and soil water potential measurements, while providing important linkages with other relevant observation networks. (ii) Fostering an inclusive and diverse collaborative environment for all scientists studying water potential in various stages of their careers. (iii) Standardizing methodologies, processing and interpretation of water potential data through the engagement of a global community of scientists, facilitated by the dissemination of standardized protocols, best practices and early career training opportunities. (iv) Facilitating the use of the PSInet database for synthesizing knowledge and addressing prominent gaps in our understanding of plants’ physiological responses to various environmental stressors. The PSInet initiative is integral to meeting the fundamental research challenge of discerning which plant species will thrive and which will be vulnerable in a world undergoing rapid warming and increasing aridification.
Gap filling time series data typically depends on linear interpolation. More recently gap filling advancements include machine learning techniques. However, none leverage advanced learning approach that uses cohort training or a neighborhood informed approach, which is described in this report. The report also describes a physics informed approach using Reduced Order Models (ROM). There are several methods to capture the nature of the detailed system in aggregated models, however there is a trade-off for these methods developed for multiple applications. These methods have specific requirements and applications that includes consideration of dynamics or covering a larger range of operating conditions, etc. The various methods of aggregation are: 1) Thevenin equivalents for downstream networks 2) Equivalent feeder representation to capture downstream network losses accurately 3) Structured reduced order models for dynamics 4) System identification-based ROM (abstract dynamical model) Methods described in items 1 and 2 above are ideal for steady-state models and useful for this application. Of these two methods, based on the data availability, the targeted application, the reduced order model that is proposed to be developed is the equivalent feeder model representation. This includes a structure of the reduced order model whose parameters can be determined by the system load and losses with the meter measurements.
This report summarizes the outcomes of the Data Center Cohort under the Department of Energy’s Technical Assistance for Digital Assurance (TADA) initiative, aimed at enhancing grid resilience through cybersecurity, supply chain risk management (SCRM), and Cyber-Informed Engineering (CIE). The cohort engaged 17 organizations across utilities, data center operators, vendors, and technology providers in three sessions combining presentations, discussions, and exercises. Key topics included AI-driven load behavior, cybersecurity vulnerabilities in UPS/BESS and cooling systems, governance gaps at utility–data center boundaries, and supply chain integrity. Five cross-cutting themes emerged: interconnection architecture vulnerabilities, fragmented governance, AI-driven stability risks, lack of regulatory frameworks, and long-term supply chain concerns. Actionable recommendations were developed, including implementing DMZ segmentation, formalizing vendor access agreements, designing AI workload limits, and advancing standards through NERC and state-level programs. These strategies aim to strengthen resilience, clarify responsibilities, and ensure secure integration of data centers into the grid.
High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.
The effectiveness of climate change impact assessments and the development of adaptation strategies depend on the availability of high-quality future weather data. However, significant gaps exist between the needs of the energy research community and the focus of the climate modeling community, primarily due to a historical lack of communication and collaboration between the two groups. Here, to address this issue, this work provides a comprehensive overview of the critical aspects involved in creating future weather data for building and energy system modeling, including emissions scenarios, general circulation models, downscaling methods, categories of future weather data, and uncertainties in climate simulations. Moreover, it critically evaluates the applicability and suitability of various types of future weather data in five key application scenarios: energy use analysis, resilience analysis, HVAC design, utility-scale analysis, and renewable energy analysis. Finally, this work presents recommendations for high-level actions and research directions to foster collaboration between the energy research and climate modeling communities and to promote the integration of future weather data into energy codes and the design practices of buildings and energy systems.
Traffic simulation is an effective tool for urban planners, traffic engineers, and researchers to study traffic. In particular, microscopic traffic simulation, which simulates individual vehicles’ movements within a transportation network, has demonstrated its importance in analyzing and managing transportation systems. However, integrating data from various sources, generating traffic scenarios, and importing information into traffic simulators to conduct microscopic simulations have always been a challenge. This paper presents a solution to overcome this challenge: RealTwin, a comprehensive tool for automated scenario generation for microscopic traffic simulation. Following a streamlined scenario generation and calibration workflow, RealTwin effectively bridges gaps between traffic data from various sources and traffic simulators, making microscopic traffic simulation more accessible for researchers and engineers across various levels of expertise. Using RealTwin to generate a real-world traffic scenario in Simulation of Urban Mobility (SUMO), VISSIM, and AIMSUN, RealTwin’s ability is demonstrated in the construction of realistic and consistent traffic scenarios in different simulators. Furthermore, this paper introduces and illustrates RealTwin’s capability for technology (e.g., autonomous vehicle) scenario generation. This feature can contribute to more comprehensive microscopic simulations, facilitating the analysis of potential effects of various technological innovations on mobility, energy efficiency, and safety. Finally, RealTwin is used to calibrate a simulation in SUMO. In conclusion, the calibration module enhances RealTwin’s ability to generate consistent simulations across different platforms and more realistic simulations that reflect real-world traffic operations.
The development of nuclear reaction models for the production of evaluated nuclear data has traditionally been performed by comparing measured cross sections with predictions from reaction model codes whose physical input parameters are adjusted to obtain the best agreement between measured and modeled results. To more directly probe reaction model inputs, this work introduces a forward modeling approach to experimental reaction cross-section determination, where the most important physical input parameters to reaction model calculations are obtained via 𝜒 2 minimization between measured and calculated observables. This was demonstrated using data collected by the Gamma Energy Neutron Energy Spectrometer for Inelastic Scattering (GENESIS) at the 88-inch cyclotron at Lawrence Berkeley National Laboratory, a detection array consisting of organic liquid scintillators and high-purity germanium (HPGe) detectors. Using a broad-spectrum neutron beam and a 99.98%-enriched 56 Fe target, GENESIS was used to perform a simultaneous measurement of 56 Fe 𝛾-ray production cross sections and secondary neutron energy and angle distributions. The results of the forward modeling approach to the determination of energy-differential 𝛾-ray production cross sections for the yrast 4 + → 2 + and 6 + → 4 + transitions, as well as eight other off-yrast transitions, were compared against those obtained using conventional techniques, and the results are in good agreement. In addition to discrete 𝛾-ray yield total scattered neutron energy-angular distributions as a function of incident neutron energy were also obtained using forward modeling and found to agree with evaluated data, with the exception of elastic scattering at small angles. The fitted reaction model parameters obtained through forward modeling were also used to calculate the cross section for the unobserved (𝑛, 2𝑛) reaction; excellent agreement with the current evaluation was obtained, providing a validation of the predictive capabilities of the forward model approach. This work bridges the gap between nuclear data experiment and evaluation by providing a new means for extracting inelastic neutron-scattering cross sections and neutron-induced 𝛾-ray production data while directly probing reaction model physics.
The commercial-scale deployment of floating offshore wind (FOW) projects is expected to take place in a diverse range of sites that may differ significantly from existing fixed-bottom projects. FOW farms are particularly sensitive to the water depth and the meteorological and oceanographic (metocean) and geotechnical conditions at the project site due to the wave-induced system motions and loads as well as the anchoring system constraints imposed by the seafloor conditions. Uncertainty around the site conditions will permeate through all aspects of project design, leading to suboptimal and overly conservative designs, increased costs, and adversely affected performance. As FOW expands into a global industry, metocean and geotechnical conditions will increasingly vary for projects located in different geographic regions or in far-from-shore, deep-water sites. This study represents the outputs of work package 1 of International Energy Agency Wind Task 49, which focuses on the integrated design of floating wind arrays. The primary goal of this study is to establish the type of parameters and constraints required to characterize FOW array reference sites; provide a realistic and publicly available set of reference site conditions to the FOW community as a baseline set of data for individual research projects; identify and categorize any critical gaps in the existing data or methodologies required to define reference site characteristics; and inform and support the design of reference FOW arrays. A building block concept was developed for synthesizing reference sites for the design of FOW arrays. The building blocks include three classes of site conditions focusing on the techno-economic design of FOW projects: metocean conditions, seabed conditions, and coastal infrastructure. All reference site data produced and collected in this study are publicly available.
The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.