Search NASA⌕ Search

SEARCH · Search NASA

Results for “data storage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

GEESS as a Mechanism to Facilitate the Commercialization of Geologic Carbon Sequestration

This is a presentation featuring an overview of the Geoanalytical Economic Evaluation of Saline Storage (GEESS) project and latest results. The GEESS project has worked towards characterizing 57 geologic saline formations targeted for geologic carbon sequestration (GCS) using publicly available datasets. The GEESS system consists of high spatial resolution datasets (up to a 5 km grid spacing) that characterize critical geologic parameters such as depth, thickness, porosity, permeability, fracture pressure, and more. Further, GEESS geologic data were exercised using the FECM/NETL Saline Storage Cost Model (CO2_S_COM) to estimate CO2 plume sizes and the first-year break-even price of CO2 at the grid point level. GEESS is now available on NETL’s Energy Data Exchange (EDX).

Eppink, Jeffrey↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Overview of the Geoanalytical Economic Evaluation of Saline Storage (GEESS) System

This report provides an overview on the Geoanalytical Economic Evaluation of Saline Storage (GEESS) system which characterized 57 geologic saline formations targeted for geologic carbon storage across the lower-48 U.S. states using publicly available resources. The GEESS system consists of high spatial resolution datasets (up to 5 km grid spacing) that provide key geologic parameters (e.g., depth, thickness, permeability), and estimates of CO 2 plume size and CO 2 first-year break-even price that were determined by exercising GEESS data through the FECM/NETL CO 2 Saline Storage Cost Model (CO2_S_COM). The GEESS geodatabase is available on NETL’s EDX at GEESS Geodatabase.

54 ENVIRONMENTAL SCIENCES↗

Solar and Battery Storage Permitting and Siting Requirements - Solar Prize Round 7 (CRADA Final Report)

The purpose of this research project was to aggregate zoning and permitting data for utility-scale solar PV and battery energy storage systems. The ultimate goal of this collaboration is to lower solar and battery energy storage system soft costs by streamlining regulatory due diligence and reducing the burden of conducting feasibility assessments for solar and storage systems. The below sections describe the specific research completed by NLR (the contractor) in furtherance of the agreement with Vanox (the participant).

14 SOLAR ENERGY↗

Machine Learning Applications in Analyzing the Role of Shale Barriers and Baffles for CO2 Storage

This study uses machine learning to analyze microseismic data from the Illinois Basin Decatur Project (IBDP) and quantify CO₂ plume extents. By leveraging well logs, microseismic records, and CO₂ injection metrics, the research predicts subsurface CO₂ plume dynamics. Findings show vertical clustering of microseismic events near the injection well, with CO₂ periodically breaching barriers due to buoyancy. K-Means clustering performed best, achieving the highest Silhouette Score and lowest Davies-Bouldin Index. This capability is crucial for real-time monitoring and management of CO₂ sequestration sites, validated against physical models and IBDP data, reinforcing CO₂ geological sequestration's viability and enhancing management tools.

Carr, Timothy↗

Myna

The additive manufacturing (AM) community has been developing digital factory tools over the past decade to better leverage the multi-modal process data coming out of the advanced manufacturing process. As a result, numerous databases of additive manufacturing process data exist in the literature and in the archival storage of disparate research groups. While some efforts have been made to create a standard ontology for storing and sharing AM data, in practice a variety of data structures are used to store AM build data, even within a single institution. This causes many problems for maintainability and extensibility when attempting to integrate computational modeling tools with experimental data to either validate models or to provide further insight into results and trends. Myna is a Python-based framework that aims to decrease the effort needed to connect individual computational models to the variety of AM process data that exist in different research groups and institutions. This type of software is sometimes referred to as "middleware" or “glueware,” in that it connects disparate databases and applications into a single computational ecosystem. Instead of maintaining unique interfaces between each application and each database, developers can create a single interface from each application to Myna and thereby gain access to the implemented database connections. Similarly, developing a database connection in Myna provides access to the developed simulation applications. This framework greatly simplifies the maintainability of model applications that rely on experimental data. Using external simulation tools, users will also be able to run pre-configured workflows using the built-in workflow manager. Several examples of input files are provided with Myna for different workflows, including melt pool geometry predictions and detailed melt pool and solidification microstructure predictions.

Knapp, GerryL. [Oak Ridge National Laboratory (ORN↗

One Earth Energy Seismic Interpretation

The objectives of the Illinois Storage Corridor (ISC) project are to accelerate commercial deployment of carbon capture utilization and storage at two individual sites and receive approvals for Underground Injection Control (UIC) Class VI permits for construction at each site (ISC Project Narrative, 2020). As part of this project, and as part of the subsurface geologic characterization, 2D seismic data was acquired at both sites. This report summarizes the findings from the 2D and 3D seismic interpretation at the One Earth Energy site near Gibson City, Illinois. The seismic data confirms the stratigraphic continuity of the Mt. Simon Arkose Zone storage interval and the Eau Claire confining unit across the project area. The seismic data also indicates that there are faults that transect the Mt. Simon Arkose Zone Sandstone storage reservoir within the modeled CO 2 plume (for more detailed information, see Faults and Fractures section of One Earth Energy Class VI Permit applications). However, the seismic data also shows that there are no faults within the modeled CO 2 plume that transect the confining unit Eau Claire Formation. The faults that transect the Mt. Simon Arkose Zone Sandstone storage reservoir all tip out in the Lower Mt. Simon Formation and do not reach the overlying Eau Claire confining unit. A small 3D survey acquired around the One Earth Energy #1 characterization well confirms these findings.

20 FOSSIL-FUELED POWER PLANTS↗

Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression

Handling large-scale scientific data in high-performance computing (HPC) environments poses significant challenges, including excessive I/O, high storage costs, and slow query performance. Traditional approaches often require full data decompression and scans, making them impractical for real-time or interactive analysis. To address these limitations, we introduce Eureka, a unified data-index co-compression framework that enables fine-grained access and efficient range queries on compressed scientific datasets. Eureka integrates spatial domain decomposition with block-wise error-bounded lossy compression to support selective decompression. It constructs a hierarchical AVL-tree index during compression to capture block-level value ranges, enabling fast pruning during query execution. To reduce metadata overhead, the index itself is also compressed while ensuring recall-preserving results. Experiments on six diverse HPC simulation datasets show that Eureka achieves up to 25x data compression and over 300x index compression, surpassing state-of-the-art compressors such as SZ3 and ZFP in rate-distortion performance. Additionally, Eureka delivers over 30x speedup for low-selectivity range queries, making it a scalable and efficient solution for modern scientific data analysis.

Yan, Ning↗

Field Test Report Neutron Scintillator Array Dry Storage Cask Scanner FY2024

During two weeks of Field Testing at the Idaho National Laboratory INTEC Cask Farm in July and August 2024, the LLNL Dry Storage Cask Scanner Array was lifted on top of an MC-10 dry storage fuel cask and operated to acquire neutron and gamma-ray data from the 24 fuel bundle positions. Neutron and gamma-ray data acquisition scans across the top of the cask of varying dwell times were performed July 15-18, 2024 and August 19-22, 2024 to evaluate the ability of the scanner data to reveal asymmetries in the fuel positions that reflect asymmetries in the MC-10 cask fuel bundle loading. The MC-10 cask 24 position fuel bundle loading at the INTEC Cask Farm is well documented, including the locations of six empty fuel bundle positions. This loading presents an opportunity to test the ability of the scanner system to detect diversion of spent fuel bundles as well as to validate the MC-10 cask MCNP modeling. The cask scanner array consists of six Stilbene crystal scintillator detectors and a linear actuator frame that moves the six detectors across the MC-10 dry storage cask to obtain data above each of the 24 fuel bundle positions. The detectors are connected to a pulse-shape discrimination data acquisition system capable of generating separate neutron and gamma-ray spectra for each detector and for each scan position. From the prior single detector Field Test in 2021 and iteration with MCNP modeling, the neutron and gamma-ray data were analyzed in multiple energy regions to identify an analysis method that would provide the strongest and most consistent signature of the asymmetric MC-10 cask fuel loading1 . From both the 2021 Field Test and the current Field Test results, the neutron capture gamma-ray count rate around 2.2 MeV provides the strongest signature of the asymmetric MC-10 cask fuel loading and has qualitative agreement with MCNP calculations. Counting all gamma-rays produces a similar signature. Neutrons emerging from the cask top are moderated and captured by the hydrogen in the polyethylene moderator and scintillator detector, producing a 2.2 MeV gamma ray which is seen in the scintillator gamma-ray spectrum. The count rate in the 2.2 MeV gamma-ray region is ~50 c/s, which is ~1000x higher than the ~0.05 n/s rate in the > 4MeV neutron region, and ~50x greater than the ~1 n/s rate in the neutrons > 500 keV region. Analysis of the 2.2 MeV neutron-capture Compton-scattered gamma-rays produces a statistically significant signature of the INTEC Cask Farm MC-10 asymmetric fuel loading. MCNP simulations indicate that the average neutron energy spectrum offers the potential to detect a large asymmetry from several missing bundles as well as individual missing fuel bundles. Testing this feature will require measurements on a cask with single missing elements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

Managing negative values is reservoir inflow computation: A case study

Reservoir inflow is conventionally estimated using the water balance method, which involves the reservoir release and the change in storage during the period considered. As a result, the estimated inflow may sometimes be negative as the errors involved in each input variable build-up to the output. In our study, the fleet data was provided by the Tennessee Valley Authority (TVA) for their Norris Hydropower facility. Unlike the flow release data, which was readily accessible, the change in storage had to be calculated using the reservoir elevation and volume relationship. The original inflow estimates produced a wide range of negative values with large outliers, making it difficult to visualize the current trends. This paper describes a methodology to remove the negative values encountered during the inflow computation, and the results were analyzed by correlating with the nearby streamflow gaging stations.

Shibu, Asha↗

Reservoir Storage Capacity Change (ResCap)

Overview Storage capacity is an essential reservoir metric that is directly linked to various water management and energy objectives. Accurate reporting and tracking of change in storage over time is crucial for the safe and reliable operation of the associated dam. While storage information is available for many reservoirs through the National Inventory of Dams, additional details, e.g. water elevation levels as well as changes over time are not included. This dataset contains reservoir storage capacities based on conducted surveys in CONUS. To represent changes in a reservoir’s storage over time, the storage capacity as determined by the first and last conducted survey is listed. The level of detail of surveys can vary greatly and improved with technological advancements. Therefore, the type of survey and year when it was conducted is noted. To ensure a fair comparison of storage capacities, the water elevation level along with the corresponding operation of the dam is reported. Structural changes, e.g. heightening of a dam will have an influence on the storage capacity and are therefore also mentioned. A total of 739 different reservoir storage capacity comparisons are listed, with some reservoirs represented more than once (storage capacity comparison at different water elevation levels). Methodology Data were acquired from USBR reservoir survey reports, TWDB lake survey reports and elevation-area-capacity tables, the RSI Web Portal and the NID (USACE, 2024). Initial storage capacity along with year and type of survey record is compared to the most recent reported storage capacity, survey type and year. Comparison elevation in feet as well as comparison elevation type were either extracted from survey reports (USBR, TWDB) or the Web Portal (RSI) and in some cases cross-referenced with data from other sources (Water Management Data, USACE, Water Data for Texas, TWDB).

Chu, Antonia [ORNL] (ORCID:0009000510540427)↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

The Role of Snowmelt and Subsurface Heterogeneity in Headwater Hydrology of a Mountainous Catchment in Colorado: A Model‐Data Integration Approach

Mountainous headwater streams are sustained by both snowmelt‐driven streamflow and groundwater discharge in the Upper Colorado River Basin. However, predicting headwater stream discharge magnitude and peak flow timing is challenging in mountainous terrains, where snowmelt rates vary with vegetation type and elevation, and heterogeneous subsurface physical properties influence groundwater storage and its release. We used a model‐data integration approach to investigate the roles of snowmelt and subsurface structure in stream discharge and groundwater level. We ran an ensemble of 100 integrated surface‐subsurface hydrologic models for a mountainous headwater catchment near Crested Butte, Colorado, USA. We also evaluated and calibrated these models against observed data sets, including snow depth measurements using distributed temperature probes, stream discharge, and groundwater levels. Calibration with multiple data sources using neural density estimators has further constrained uncertainty in subsurface properties and snowmelt rates. Results indicated that observed slower snowmelt rates in evergreen forests delayed the peak flow and baseflow onset. In upstream areas with lower subsurface permeability, water was stored within the subsurface but was not released as interflow or shallow groundwater flow, and thereby not contributing to downstream streamflow during recession limb periods. Double peaks in groundwater occurred in areas with spatial subsurface heterogeneity, in our case due to the contrast between granodiorite and Mancos shale. These process‐based insights into groundwater and snowmelt dynamics in mountainous headwaters will help improve predictions of headwater hydrology.

Wang, Lijing [University of Connecticut, Storrs, C↗

Learning from Arctic Microgrids: Cost and Resiliency Projections for Renewable Energy Expansion with Hydrogen and Battery Storage

Electricity in rural Alaska is provided by more than 200 standalone microgrid systems powered predominantly by diesel generators. Incorporating renewable energy generation and storage to these systems can reduce their reliance on costly imported fuel and improve sustainability; however, uncertainty remains about optimal grid architectures to minimize cost, including how and when to incorporate long-duration energy storage. This study implements a novel, multi-pronged approach to assess the techno-economic feasibility of future energy pathways in the community of Kotzebue, which has already successfully deployed solar photovoltaics, wind turbines, and battery storage systems. Using real community load, resource, and generation data, we develop a series of comparison models using the HOMER Pro software tool to evaluate microgrid architectures to meet over 90% of the annual community electricity demand with renewable generation, considering both battery and hydrogen energy storage. We find that near-term planned capacity expansions in the community could enable over 50% renewable generation and reduce the total cost of energy. Additional build-outs to reach 75% renewable generation are shown to be competitive with current costs, but further capacity expansion is not currently economical. We additionally include a cost sensitivity analysis and a storage capacity sizing assessment that suggest hydrogen storage may be economically viable if battery costs increase, but large-scale seasonal storage via hydrogen is currently unlikely to be cost-effective nor practical for the region considered. While these findings are based on data and community priorities in Kotzebue, we expect this approach to be relevant to many communities in the Arctic and Sub-Arctic regions working to improve energy reliability, sustainability, and security.

25 ENERGY STORAGE↗