Search NASA⌕ Search

SEARCH · Search NASA

Results for “DATA STORAGE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

LANL Meteorological Program: 2023 Data Completeness/Quality Report

Los Alamos National Laboratory (LANL) operates seven mesa-top instrumented meteorology towers: Technical Area (TA) 6, TA-49, TA-53, TA-54, TA-63, TA-54B, and TA-16B. An additional instrumented tower is located in Mortandad Canyon (TA-5 MDCN), and there is a rain gauge at North Community (NCOM), located within the town of Los Alamos. The 10 meter (m) towers at TA-63, TA-54B, and TA-16B have been in testing since they were installed in 2021, and will be included in a future data completeness report. A description of the meteorology monitoring network, prior to the installation of the TA-63, TA-54B, and TA-16B is found in Dewart and Boggs (2014). Four of the mesa-top towers (e.g., TA-6, TA-49, TA-53, and TA-54) are instrumented at the 1.2 m, 11.5 m, 23 m, and 46 m levels. In addition, the TA-6 tower is instrumented at the 92 m level. The TA-5 MDCN tower is 10 m in height and is instrumented at 1.2 m and 10 m. Data are collected and averaged every 15 minutes. Range checking is performed on each measurement every 15 minutes; data that are beyond normal ranges are eliminated from the data set and replaced by a code for missing data. In addition, data are reviewed weekly by qualified meteorologists to identify bad data not identified by the range checking technique. The data steward eliminates these data from the data set and replaces them with a code for missing data. The instrument technicians also review that data and schedule instrument replacement, as required. All instruments are calibrated at a frequency that meets the criteria identified in ANSI/ANS-3.11-2015. Data completeness is determined by the number of total 15-minute records available versus the number of possible measurements for the entire year. As a rule, the meteorologists do not attempt to estimate data that are eliminated as bad data. Original datalogger records, including bad data, can be recalled from program archival storage.

54 ENVIRONMENTAL SCIENCES↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

A Novel and Scalable Method for Microencapsulating Salt Hydrate Phase Change Materials in Core–Shell Fibers

Phase change materials (PCMs) are in high demand for applications such as thermal energy storage in buildings, electronics cooling, and thermal management of electric vehicle batteries and data centers. Among these materials, salt hydrate PCMs are particularly attractive due to their high thermal energy storage capacity and low cost. However, they suffer from two major issues: leakage in the melted phase and phase segregation during phase transitions. Microencapsulation is the primary process capable of addressing both of these challenges. However, there is no reliable or scalable method available for microencapsulating salt hydrate PCMs. As a result, the full potential of salt hydrates for building and data center applications has yet to be realized. In this work, we present an innovative method for the microencapsulation of salt hydrate PCMs using a co‐axial pushing technique. This process creates core–shell fibers, with the salt hydrate as the core and a polymer as the shell. Our approach demonstrates strong potential for scalable microencapsulation of salt hydrate PCMs. In conclusion, achieving scalability could enable their widespread use in applications such as data center cooling, battery thermal management, and building climate control.

Sharma, Jaswinder [Oak Ridge National Laboratory (↗

Prospecting for Critical Minerals and Rare Earth Elements from Marcellus Shale in the Western Portion of the Appalachian Basin with Non-Destructive Core Characterization

Identification of sources for domestic critical minerals and rare earth elements (CM/REE) has been deemed essential for the energy transition by the United States Department of Energy (DOE). The U.S. DOE’s National Energy Technology Laboratory’s (NETL) Geomaterials Characterization Laboratory has performed non-destructive core characterizations on energy-relevant rock cores for the past decade. During this time, NETL has published over 36 technical reports and made the associated data publicly available. Much of this work focuses on unconventional shale gas, subsurface carbon storage systems, and carbon-ore. These efforts provide cm-scale petrophysical and elemental data, photographic documentation, detailed core descriptions, and computed tomography (CT) data for each well. This provides a first phase prospecting resource for CM/REE resources and can provide a map for pin-pointing intervals and lithologies for further development. Using historical core characterization data from 12 Marcellus wells from the western portion of the Appalachian Basin, this study builds an improved understanding of the chemostratigraphy of the basin. X-ray fluorescence (XRF) and CT images were used to determine lithologic intervals and potential ore bodies for further analysis, including benchtop digestion and inductively coupled plasma mass spectrometry (ICP-MS) to better understand the CM/REE enrichments.

Paronish, Thomas J.↗

What to Support When You’re Compressing

Over the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research.

Error-Bounded Lossy Compression↗

A Convolution Neural Network for Voltage Event Classification at a Photovoltaic Inverter

This paper presents a convolutional neural network (CNN) developed to identify voltage events in photovoltaic (PV) inverters. The CNN is trained on synthetic data generated using the IEEE 13-bus distribution feeder model and evaluated on field measured data collected from Energy Northwest’s Horn Rapids Solar, Storage, and Training (HRSST) facility. The study focuses on two common voltage events: faults and voltage sags. The CNN is configured to analyze voltage and current waveforms from three-phase PV systems, demonstrating excellent accuracy during training. Field data from the HRSST facility is employed to assess its real-world performance, where the CNN achieves perfect identification of faults and voltage sags in a sample of nine events. This work highlights the potential of the proposed method to enhance PV protection schemes, providing a robust foundation for improved voltage event detection and grid reliability.

Cornachione, Matthew A.↗

Immobilization of Urease for continuous flow conversion of waste urea

An efficient and robust system for the urease catalyzed conversion of urea to ammonia has been developed using urease immobilized on modified agarose beads. Two different immobilization strategies, adsorption and covalent binding were studies using six different types of modified agarose beads. The immobilization of urease on each of the beads was studied at different concentrations and times using immobilization efficiency as a measure of success. The data from these experiments was used to identify potential candidates for immobilization scale up and implementation into the continuous flow system. The enzyme was then immobilized on the candidates in a packed bed reactor and the optimal flow rate and storage stability was determined. Future work will utilize the data obtained from these experiments to expand to other resins and enzymes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Artificial Scientist: in-Transit Machine Learning of Plasma Simulations

Large-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-incell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list.

Kelling, Jeffrey [Helmholtz-Zentrum Dresden Rossen↗

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE↗

Basin-Scale Structural Features Database: Spatial Datasets to Support Carbon Storage Resource Assessments

Presentation slides on "Basin-Scale Structural Features Database: Spatial Datasets to Support Carbon Storage Resource Assessments" for CCUS 2025 Annual Meeting. The Basin-Scale Structural Features database contains a series of basin-scale spatial datasets representing structural features, including faults, fractures, folds, and earthquakes. Designed to support carbon storage feasibility and resources assessments for Carbon Capture and Storage (CCS) projects, the database leverages publicly available data resources from authoritative sources (e.g. US Geological Survey, State Geologic Surveys), and aims to help users better understand basin-scale structural features, as well as potential data gaps in areas with sparse information.

basin scale↗

Site Characterization of the Highest-Priority Geologic Formations for CO2 Storage in Wyoming

The project Site Characterization of the Highest-Priority Geologic Formations for CO2 Storage in Wyoming is one of 9 site characterization projects that were implemented as part of ARRA (American Recovery and Reinvestment Act). Data from this project was used to improve resolution of data in NATCARB in the area of study. Data related to this study has already been incorporated in NATCARB Atlas. The Wyoming Carbon Underground Storage Project (WY-CUSP) consisted of CO2 storage site characterization and evaluation, focusing on Wyoming’s most promising CO2 storage reservoirs (the Pennsylvanian Weber/Tensleep Sandstone and Mississippian Madison Limestone) and premier CO2 storage site (Rock Springs Uplift). Results from the WY-CUSP project suggest the two reservoirs could store up to 17,000 million tons of CO2. The WY-CUSP team drilled a stratigraphic test well and acquired a 3-D seismic survey covering 25 square miles of the Rock Springs Uplift site. The team retrieved 916 feet of core from the 12,810-foot-deep well, along with a complete log suite, borehole images, fluid samples, and other data. Project partners (1) provided continuous visual documentation of the core, including grain size, mineralogy, facies distribution, and porosity; (2) performed continuous permeability and velocity scans of selected reservoir intervals; and (3) chemically analyzed the fluid samples. WY-CUSP scientists integrated seismic attributes with observations from log suites, a VSP survey, core, fluid samples, and laboratory analyses, including continuous permeability scans. From these integrations, researchers constructed 3-D spatial distribution volumes of reservoir and seal properties that represent geological heterogeneity at the targeted CO2 storage site. The WY-CUSP team used this data to perform new CO2 plume migration simulations. Baker Hughes, Inc., completed a series of small-scale, in-situ water injectivity measurements. A database was formed when observations, analyses, and experiments from the stratigraphic test well were integrated. Correlation of these data allowed petrophysical parameters to be extrapolated from the test well out into the storage domain (5x5 mile 3-D seismic survey volume). This resulted in an improved, realistic understanding of performance assessments for potential CO2 storage scenarios. The WY-CUSP team worked on (1) improving CO2 storage resource estimates, (2) establishing long-term integrity and permanence of confining layers, (3) designing a profitable strategy for pressure management, and (4) evaluating the utilization of stored CO2 at the Rock Spring Uplift. Finally, Baker Hughes developed a microseismic baseline for the test site using in-bore geophones to complete field operations.

3-D seismic↗

DaYu: Optimizing Distributed Scientific Workflows by Decoding Dataflow Semantics and Dynamics

The combination of ever-growing scientific datasets and distributed workflow complexity creates I/O performance bottlenecks due to data volume, velocity, and variety. Although the increasing use of descriptive data formats (e.g., HDF5, netCDF) helps organize these datasets, it also creates obscure bottlenecks due to the need to translate high level operations into file addresses and then into low-level I/O operations. To address this challenge, we introduce DaYu, a method and toolset for analyzing (a) semantic relationships between logical datasets and file addresses, (b) how dataset operations translate into I/O, and (c) the combination across entire workflows. DaYu's analysis and visualization enables identification of critical bottlenecks and reasoning about remediation. We describe our methodology and propose optimization guidelines. Evaluation on scientific workflows demonstrates up to 3.7x performance improvements in I/O time for obscure bottlenecks. The time and storage overhead for DaYu's time-ordered data is typically under 0.2% of runtime and 0.25% of data volume, respectively.

Tang, Meng↗

Continuing Development of the Nuclear Data Processing Code AMPX [Poster]

The ENDF/B-VIII.1 evaluation library has seen a great growth in the thermal neutron scattering sub-library. The SCALE code system has traditionally approached CE transport by assuming that the CE library on disk represented the fully expanded cumulative probability distributions, conditional on exiting angle and marginal on exiting energy. While this is a complete description of the data, it comes at the potential cost of large amounts of on-disk storage. This approach was strained by several TSL files in ENDF/B-VIII.1, such as graphite, which contained data for a large number of Bragg edges. In the fully expanded probability distributions, this was found to be a disproportionately large fraction of the SCALE CE library.

GNDS↗

Assessing hydrogen supply chains: An integrated review of leakage and energy efficiency studies

This paper examines hydrogen leakage and efficiency across the supply chain for liquid, gaseous, and mixed hydrogen systems. These factors are crucial for assessing hydrogen's role in mitigating emissions and facilitating a clean energy transition. Drawing on a comprehensive review of existing literature and model-based analysis, the study compiles leakage rates and efficiency metrics at each stage of the supply chain: production, storage, transmission, distribution, and end-use. These data inform system scenarios that estimate the impact of leakage on overall performance and climate benefits. The analysis also identifies persistent data gaps, particularly for liquid and mixed system configurations, and outlines priorities for future research. A comparison of hydrogen system types shows that gaseous pathways generally achieve the highest efficiencies (28 %–39 %) and the lowest leakage rates (∼4.5 %) across the supply chain. Liquid hydrogen systems, while favorable for long-distance and high-volume transport due to their higher energy density, exhibit lower efficiency (∼28 %) and a greater leakage potential (∼12 %). Mixed systems, which combine gaseous and liquid elements (e.g., pipeline transmission followed by liquefaction and truck distribution), show compounded energy losses and moderate-to-high leakage rates (6.8 %–9.4 %), highlighting trade-offs associated with added system complexity. The study highlights opportunities for technological advancements, including optimizing liquefaction, enhancing insulation for storage and transportation, and refining refueling equipment. These improvements are crucial for maximizing the climate benefits of hydrogen. The results offer actionable insights for researchers, industry, and policymakers working to develop low-leakage, high-efficiency hydrogen infrastructure.

08 HYDROGEN↗

Repository of HydroSMADE: Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites

This repository presents HydroSMADE—Hydropower Site-level Monthly Availability Data Ensemble, a new open dataset that provides monthly hydropower availability for 1,593 existing and 124,333 potential sites worldwide over the period 1950–2100. The dataset is generated by using a global hydrologic model (Xanthos) with explicit representation of hydropower operation. Specifically, HydroSMADE distinguishes between storage and diversion sites, applies optimized operating rules, and incorporates site-specific characteristics such as generation capacity, maximum turbine flow, and reservoir storage. Driven by bias-corrected meteorological inputs, the data is provided for 30 alternative future scenarios. The scenarios consist of the full factorial combination of three standard CMIP6 atmospheric forcing pathways (SSP1-2.6, SSP3-7.0, and SSP5-8.5) and ten CMIP6 General Circulation Models (GCMs): GFDL-ESM4, IPSL-CM6A-LR, MPI-ESM1-2-HR, MRI-ESM2-0, EC-Earth3, CanESM5, MIROC6, CNRM-ESM2-1, UKESM1-0-LL, and CNRM-CM6-1. The repository contains a total of 122 files: a text file (readme.txt) containing a brief description of the included data, a CSV file containing site attributes, and the remaining 120 files (in CSV) containing site-level monthly hydropower availability. Example Jupyter Notebooks to explore the HydroSMADE dataset are available on GitHub at https://github.com/kamal0013/HydroSMADE More details on the methods and technical validation of HydroSMADE are available in the following paper by the same authors: Chowdhury, A. K., Abeshu, G. W., Zhao, M., Wild, T. B., Hassan, N., Ying, Z., Kim, G. J., Matthew, B., Jonathan, L., & Li, H.-Y. (Submitted). Hydropower Site-level Monthly Availability Data Ensemble for 1950-2100 at Existing and Potential Global Sites.

Existing and Potential Sites↗

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Conference presentation at Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM↗

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Extended abstract for Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM↗