Search NASA⌕ Search

SEARCH · Search NASA

Results for “data storage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

Extending Rucio with modern cloud storage support

Rucio is a software framework designed to facilitate scientific collaborations in efficiently organising, managing, and accessing extensive volumes of data through customizable policies. The framework enables data distribution across globally distributed locations and heterogeneous data centres, integrating various storage and network technologies into a unified federated entity. Rucio offers advanced features like distributed data recovery and adaptive replication, and it exhibits high scalability, modularity, and extensibility. Originally developed to meet the requirements of the high-energy physics experiment ATLAS, Rucio has been continuously expanded to support LHC experiments and diverse scientific communities. Recent R&D projects within these communities have evaluated the integration of both private and commercially-provided cloud storage systems, leading to the development of additional functionalities for seamless integration within Rucio. Furthermore, the underlying systems, FTS and GFAL/Davix, have been extended to cater to specific use cases. This contribution focuses on the technical aspects of this work, particularly the challenges encountered in building a generic interface for self-hosted cloud storage, such as MinIO or CEPH S3 Gateway, and established providers like Google Cloud Storage and Amazon Simple Storage Service. Additionally, the integration of decentralised clouds like SEAL is explored. Key aspects, including authentication and authorisation, direct and remote access, throughput and cost estimation, are highlighted, along with shared experiences in daily operations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

wmpy-power: A Python package for process-based regional hydropower simulation

Hydropower is an important source of renewable energy in many parts of the world. The generation potential for a hydropower facility can vary greatly due to fluctuations in precipitation and snowmelt patterns impacting streamflow and reservoir storage. Human activities such as irrigation, manufacturing, and hydration can also influence water availability at nearby and downstream facilities. wmpy-power--the hydropower model described in this work--is process-based, leveraging explicit reservoir storage and release data to address impacts on hydropower from climate change and human adaptive behaviors to inform long-term planning and resource-adequacy considerations.

13 HYDRO ENERGY↗

Demand response of loads having thermal reserves

Systems and methods are described herein that improve grid performance by smoothing demand using thermal reserves. The smoothed demand can reduce peak loads as well as the ramp rate of demand that will otherwise require the use of inefficient, expensive generation sources. These improvements are tied to the selective switching on or off electrical loads that are coupled to thermal reserves, effectively using the thermal reserves as an energy storage mechanism. Historical data of past usage can be used to create load model and ensure that effects on customer comfort are minimized while still accomplishing the beneficial effects for the overall grid, which enables grid owners to both reduce their operational cost by avoiding expensive generation and improve system reliability by achieving more predictable power demand.

Ren, Wei↗

3D Deep Learning Joint Inversion of Active Seismic Full Waveform and Passive Seismic Traveltime Data for Reservoir Imaging and Uncertainty Quantification

Here, we present deep learning (DL) networks for three-dimensional (3D) joint inversion of active seismic full waveform and passive seismic traveltime data to image reservoirs and their properties and quantify imaging uncertainties. Active seismic full-waveform data can provide high-resolution monitoring images but are collected only intermittently because of their high acquisition cost. In contrast, passive seismic data can be gathered at relatively low cost between regular active surveys, although their imaging quality can be compromised by factors such as low signal-to-noise ratios and limited ray coverage of the target. Although these datasets are routinely acquired together at CO 2 storage sites, their combined inversion within a 3D DL framework has not been previously demonstrated. To our knowledge, this is the first study to address this gap, combining the strength of both data types. For efficient data storage and DL training with large 3D seismic datasets, we use a 3D data matrix in which a random number of passive seismic traveltime data are stored as parabolic envelopes using one-hot encoding and a 3D full-waveform data matrix in which multiple shot gathers are summed. Two network architectures are evaluated: a single-encoder U-Net for single-data type inversion and a dual-encoder U-Net for joint inversion of active and passive seismic data. We also evaluate the single-encoder U-Net for joint inversion by concatenating full-waveform data and traveltime data. We propose a systematic approach for selecting an optimal dropout rate that balances regularization during training and Monte Carlo dropout-based uncertainty quantification during prediction by examining the correlation coefficient between standard deviation and prediction error, along with the training misfit, across a range of dropout rates. 3D DL inversion experiments include five different network configurations, with evaluations under ideal, noisy and dropout-enabled conditions. Both model and data uncertainties are assessed, as well as their combined effects. Across all conditions, the networks consistently predict accurate CO 2 saturation models with low prediction errors, such as a structural similarity index of 0.993 and CO 2 difference of 1.1%. Uncertainty estimates show strong spatial correlation with prediction errors, confirming the effectiveness of the proposed dropout selection approach. The results demonstrate that our DL approach, utilizing compact data representations and appropriate uncertainty quantification, yields accurate subsurface images under various inversion conditions and provides valuable insights into the reliability of predictions.

Um, Evan Schankee [Lawrence Berkeley National Labo↗

Scientific Data Compression for Large Scale Computational Fluid Dynamics (CFD) Simulations

This Cooperative Research and Development Agreement (CRADA) between Oak Ridge National Laboratory (ORNL) and General Electric (GE) investigated methods for reducing the size of large computational fluid dynamics (CFD) simulation datasets using scientific data compression techniques. The work focused on adapting the MultiGrid Adaptive Reduction of Data (MGARD) compression framework and integrating it with high-performance I/O and visualization tools used in CFD workflows. MGARD uses hierarchical multilevel decomposition to enable error-controlled compression of floating-point scientific data while preserving quantities of interest. During the project, MGARD compression was integrated with the ADIOS I/O framework and visualization tools such as ParaView to enable efficient storage, transfer, and analysis of simulation data. The collaboration also explored approaches for improving compression performance for CFD data defined on unstructured meshes. Results demonstrate that scientific data compression can significantly reduce storage requirements and improve data management for large-scale CFD simulations.

97 MATHEMATICS AND COMPUTING↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Systems, methods, and devices for failure detection of one or more energy storage devices

An energy storage device management system can include a management portion for charging/discharging an energy storage device and an ultrasound interrogation portion for passing ultrasound energy through the energy storage device during charge/discharge cycles. A memory stores a stream of capture data instances derived from ultrasound energy exiting the energy storage device and baseline ultrasound data instances corresponding with the energy storage device during normal charging/discharging thereof. A processor can compare each capture data instance with the baseline ultrasound data and detect abnormal operating states of the energy storage device. A warning system can issue a notification when abnormal operating states are detected.

Kowalski, Jeffrey A.↗

On-Demand Column Joining for High Energy Physics

As the Large Hadron Collider (LHC) transitions into the High-Luminosity LHC (HL-LHC) era, the volume of data to be processed is expected to increase significantly. The CMS Experiment currently utilizes various data formats, including AOD, MiniAOD, and NanoAOD, each with different levels of detail and storage requirements. This paper addresses the challenges of data duplication and storage inefficiencies in high-energy physics (HEP) analyses by proposing an on-demand column-joining solution. This approach aims to reduce data duplication by enabling the dynamic combination of NanoAOD data with auxiliary information from larger data tiers, such as MiniAOD. The proposed solution leverages Trino, a high-performance distributed SQL query engine, to perform efficient and scalable data joins. Benchmarks using CMS OpenData demonstrate the feasibility of this approach, showing that it can handle large datasets with low latency. Integration with the scikit-hep ecosystem and the coffea analysis framework is also discussed, highlighting the potential for seamless end-to-end data processing and analysis. Ongoing and future work focuses on expanding benchmarks, integrating ServiceX for data transformation, and exploring the use of native object storage solutions.

Manganelli, Nicholas [Northeastern U.]↗

Mitigating Data Center Impact on Grid Stability: A Coordinated Control Strategy Using Verrus StabiliGrid Architecture

Large data centers, which now represent a significant and growing share of the total U.S. grid load, can inadvertently destabilize the electrical grid when they disconnect simultaneously during brief voltage disturbances. The July 10, 2024, Eastern Interconnection incident, in which a sub-100-millisecond transmission fault triggered the cascading loss of approximately 1,500 MW of data center load, illustrates this vulnerability. While commercial battery energy storage systems (BESS) deployed in data centers provide device-level fault ride-through per IEEE 1547, they lack coordination with facility protection logic and uninterruptible power supplies (UPS), limiting their effectiveness as grid-stabilizing assets. This report presents the Verrus StabiliGrid architecture, a coordinated control framework that integrates BESS, UPS, and point-of-interconnection (POI) protection settings to enable data centers to ride through both undervoltage and overvoltage grid contingencies without disconnecting. The four-step strategy encompasses: (1) high-resolution power quality monitoring to detect the grid state during events such as undervoltage, overvoltage, underfrequency, and overfrequency; (2) POI protection settings that allow for extended ride-through and grid-connected operation during grid contingencies; (3) grid state-driven autonomous dispatch of assets to improve grid resilience by reducing power draw during undervoltage or absorbing more power during overvoltage events; and (4) coordinated post-recovery dispatch of data center assets to restore firm load to pre-contingency levels. Validated through controller-hardware-in-the-loop (C-HIL) simulations at the National Laboratory of the Rockies, results show grid import restoration to pre-fault levels within 100 milliseconds of voltage recovery. This work advances the ability of data centers to transition from passive, disturbance-sensitive loads to active participants in grid stability, a capability increasingly required by emerging NERC and ERCOT regulatory frameworks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data Management in the Continuum: Cross-facility Object-based Data Transfers

Scientific workflows are evolving from relying on a monolithic storage subsystem at a single High-Performance Computing (HPC) facility to using geographically distributed file systems, repositories, and cloud storage. As a result, storing, accessing, transferring, and managing scientific data have become highly complex and prone to performance inefficiencies. This paper delves into these challenges by exploring an optimized end-to-end interface designed to seamlessly connect various local and remote storage systems, enabling efficient data movement of objects across HPC–Cloud and HPC–HPC environments. We showcase this capability through an object-focused data management runtime system, discuss the effects of relaxed consistency semantics in distributed object scenarios, and illustrate its application in an earthquake simulation workflow. Besides reducing the amount of data by selectively transferring regions of interest, our facility-local results achieved a speedup of 45 × over an optimized HDF5 usage and 15 × over the HDF5 with caching by using the new interface in PDC-XF.

Bez, Jean Luca↗

Scientific Data Management Beyond Traditional Computing Boundaries

Scientific data management is undergoing a fundamental transformation driven by the convergence of artificial intelligence (AI)/machine learning workflows, distributed computing and storage environments, and exponential data growth. Here, we analyze how these developments address current limitations while enabling new capabilities for cross-facility collaboration and AI-driven research.

Widener, Patrick [Oak Ridge National Laboratory (O↗

Uinta Basin CarbonSAFE II: Storage Complex Feasibility (Final Report)

The primary objective of this CarbonSAFE Phase II project was to establish the technical and commercial feasibility of a commercial-scale CO 2 geological storage complex for Deseret Power Electric Cooperative Bonanza Power Plant and other CO 2 sources in the northeast Uinta Basin, Utah, with the goal to securely store at least 50 million metric tons of captured CO 2 and accelerate CO 2 capture, utilization, and storage (CCUS) deployment. The project team established high-potential technical and commercial feasibility for a storage site within the east Uinta Basin (Utah), in the Cretaceous sandstones (Frontier, Dakota, and Buckhorn), Entrada Sandstone, Nugget Sandstone, and/or Weber Sandstone southwest of the Bonanza coal-fired power plant. This project collected and analyzed state-of-the-art data to characterize the storage complex consistent with Environmental Protection Agency (EPA) permitting standards. The team conducted extensive analog studies, outcrop mapping, and data sampling, which largely contributed to understanding the subsurface lithology and facies. Existing data were obtained and assessed from Utah Division of Oil, Gas, and Mining (DOGM), Utah Geological Survey (UGS), Colorado Geological Survey (CGS), U.S. Geological Survey (USGS), and EPA. These data were analyzed using state-of-the-art CCUS technologies for Societal Considerations, Site Characterization, Modeling and Simulations, Risk Assessment, Management and Monitoring, potential Underground Injection Control (UIC) Class VI Well Permitting, and Technical/Economic Feasibility. Through these high-resolution data collection and feasibility studies, this project was expected to provide a reference for initiating Underground Injection Control (UIC) and other commercial-scale geological storage permitting processes in the Western United States, ultimately contributing to the nation's decarbonization goals through low-risk, cost-effective commercial-scale carbon capture, utilization, and storage (CCUS) projects.

42 ENGINEERING↗

Numerical Investigation of High Delta T Sensible Storage Integrated CO2 Heat Pump: Preprint

To assist building heating electrification, this paper numerically investigates a load flexible heat pump system for commercial buildings. The system consists of a CO2 vapor compression cycle, a sensible thermal storage tank, and an air handling unit. The thermal storage medium is inexpensive, non-toxic and stable anti-freeze solution (30% potassium acetate). The air handing unit has an indoor coil and a ventilation coil. The system can be used to manage building electric load. During peak hours, the heat pump is off and the hot solution water is discharged from the tank to heat up the indoor air and ventilation air. During the hour of charge, the heat pump delivers hot solution water to the tank and to the air. The tank can also stand by while the heat pump provides space heating directly. We selected a medium sized office building located in Minnesota as the representative building and used EnergyPlus to obtain its 24 hour load data. We designed three storage tank volumes assuming 50 degrees C, 65 degrees C and 80 degrees C tank temperatures to independently provide the building load for 4 hours in the morning. The higher the tank temperature, the smaller the required volume, and thus higher energy density. The effective energy density is 78 with an 80 degrees C tank, and 40 kWhth/m3 with 50 degrees C. We simulated the tank integrated heat pump performance subjected to the 24-hour building load profile and ambient data. The baseline is the same system without storage tank. There was a trade-off between the storage energy density and the charging COP. The charge hour COP was 2.77 to charge the tank to 80 degrees C, and 3.01 to 50 degrees C. The proposed system could shift building load from the peak hours (8:00 - 12:00) to off-business hour (23:00 - 7:00+1). It eliminated 100% compressor electricity use during the peak hours, and avoided a peak electric power of 34 kW. The 65 degrees C tank saved 9.5 kWhe (4%) considering all day operation, which was the best balance between energy density and the system operation efficiency among the three options.

CO2 heat pump↗

Pumped Storage Hydropower Potential and Opportunities

Pumped storage hydropower (PSH) is a flexible energy storage technology with the potential to improve grid reliability, resiliency, and stability in the electric grid of the future. NREL has developed a range of data and tools to help understand opportunities for new PSH deployment, including nationwide resource assessment data, a bottom-up component-level cost model, and a lifecycle greenhouse gas emissions calculator. These datasets can then be used to inform grid planning models, analysis, and decision making to understand the role PSH can play in the power sector.

cost↗