Search NASA⌕ Search

SEARCH · Search NASA

Results for “storage throughput”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Machine learning-assisted design of metal–organic frameworks for hydrogen storage: A high-throughput screening and experimental approach

Various theoretical approaches, including big data and high-throughput screening techniques, have been explored in developing new materials due to their significant potential time-saving advantages. However, it remains a significant challenge to experimentally realize new materials that are predicted. In this study, we propose a novel materials design strategy that utilizes machine-learning (ML) techniques to predict new porous materials that show promise for hydrogen storage and are likely to be feasible to synthesize. By leveraging ML techniques and metal–organic framework (MOF) databases, we are able to predict the synthesizability of MOF structures. This is evidenced by the successful synthesis of a new vanadium-based MOF that exhibits excellent performance for cryogenic H 2 storage. Notably, the total gravimetric and volumetric H 2 uptakes are as high as 9.0 wt% and 50.0 g/L at 77 K and 150 bar. This ML-assisted materials design offers an efficient and promising approach for developing hydrogen storage materials.

08 HYDROGEN↗

When to use rsync

We have endeavored to show, using a series of data transfer results obtained from two testbeds, when to use the popular data copying tool rsync and related tools. Tests have been conducted in local area network (LAN) and wide area network (WAN) environments. We conclude that for files in a certain size range and network latency ≦ 10 ms round trip time (RTT), rsync is still useful for data moving tasks in the category 4 of the U.S. DOE Technical Report “Data Movement Categories”. For more demanding data movement requirements, tools of different classes are suggested. Sample histograms from two DOE user facilities are provided to further support our conclusions.

97 MATHEMATICS AND COMPUTING↗

dCache project status and update

The dCache project delivers an open-source, massively scalable, distributed storage system deployed internationally to satisfy today’s scientists’ ever-demanding storage requirements. Its multifaceted approach supports different use cases with the same storage, from high throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and longterm data persistence on tertiary storage. Even though dCache was initially developed for HEP experiments, today, it is used by various scientific communities, including astrophysics, biomed, and life science, each with their specific requirements. To match the needs of these new communities and keep up with the scaling demands of existing experiments, dCache is permanently evolving. With this contribution, we would like to highlight the recent developments in dCache regarding integration with CERN Tape Archive (CTA), advanced metadata handling, token-based authorization support, bulk API for QoS transitions, REST API to control interaction with the tape system, and future development directions.

Mkrtchyan, Tigran [DESY]↗

dCache: Inter-disciplinary storage system

The dCache project provides open-source software deployed internationally to satisfy ever more demanding storage requirements. Its multifaceted approach provides an integrated way of supporting different use-cases with the same storage, from high throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters and long term data persistence on a tertiary storage. Though it was originally developed for the HEP experiments, today it is used by various scientific communities, including astrophysics, biomed, life science, which have their specific requirements. In this paper we describe some of the new requirements as well as demonstrate how dCache developers are addressing them.

Mkrtchyan, Tigran↗

Variable rate neural compression for sparse detector data

Particle colliders produce data at extraordinary rates, posing major challenges for transmission and storage. High-throughput compression algorithms are therefore essential. In the sPHENIX experiment taking data at the Relativistic Heavy Ion Collider, a time projection chamber records three-dimensional (3D) particle trajectories that are highly sparse, making conventional learning-free lossy compression ineffective. Convolutional neural networks have surpassed traditional methods in compression ratio and accuracy. However, they fail to exploit sparsity for efficiency. To address these gaps, we present BCAE-VS, a bicephalous convolutional autoencoder with variable compression ratio for sparse data, which adapts compression to input complexity through key-point identification and sparse convolution. BCAE-VS achieves higher accuracy and compression ratios than prior neural approaches while being orders of magnitude smaller. Moreover, its throughput increases with sparsity—a property not observed in other methods. Although it was developed for collider experiments, BCAE-VS readily extends to other sparse data domains, such as light detection and ranging (LiDAR) sensing and 3D microscopy.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Geomancy: Automated Performance Enhancement through Data Layout Optimization

Large distributed storage systems such as high- performance computing (HPC) systems used by national or international laboratories require sufficient performance and scale for demanding scientific workloads and must handle shifting workloads with ease. Ideally, data is placed in locations to optimize performance, but the size and complexity of large storage systems inhibit rapid effective restructuring of data layouts to maintain performance as workloads shift. To address these issues, we have developed Geomancy, a tool that models the placement of data within a distributed storage system and reacts to drops in performance. Using a combination of machine learning techniques suitable for temporal modeling, Geomancy determines when and where a bottleneck may happen due to changing workloads and suggests changes in the layout that mitigate or prevent them. Our approach to optimizing throughput offers benefits for storage systems such as avoiding potential bottlenecks and increasing overall I/O throughput from 11% to 30%.

Bel, Oceane M.↗

Fast 2D Bicephalous Convolutional Autoencoder for Compressing 3D Time Projection Chamber Data

High-energy large-scale particle colliders produce data at high speed in the order of 1 terabytes per second in nuclear physics and petabytes per second in high energy physics. Developing real-time data compression algorithms to reduce such data at high throughput to fit permanent storage has drawn increasing attention. Specifically, at the newly constructed sPHENIX experiment at the Relativistic Heavy Ion Collider (RHIC), a time projection chamber is used as the main tracking detector, which records particle trajectories in a volume of three-dimensional (3D) cylinder. The resulting data are usually very sparse with occupancy around 10.8%. Such sparsity presents a challenge to conventional learning-free lossy compression algorithms, such as SZ, ZFP, and MGARD. The 3D convolutional neural network (CNN)-based approach, Bicephalous Convolutional Autoencoder (BCAE), outperforms traditional methods both in compression rate and reconstruction accuracy. BCAE can also utilize the computation power of graphical processing units suitable for deployment in a modern heterogeneous highperformance computing environment. This work introduces two BCAE variants: BCAE++ and BCAE-2D. BCAE++ achieves a 15% better compression ratio and a 77% better reconstruction accuracy measured in mean absolute error compared with BCAE. BCAE-2D treats the radial direction as the channel dimension of an image, resulting in a 3× speedup in compression throughput. In addition, we demonstrate an unbalanced autoencoder with a larger decoder can improve reconstruction accuracy without significantly sacrificing throughput. Lastly, we observe both the BCAE++ and BCAE-2D can benefit more from using half-precision mode in throughput (76 - 79% increase) without loss in reconstruction accuracy. The source code and links to data and pretrained models can be found at https://github.com/BNL-DAQ-LDRD/NeuralCompression_v2

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Computationally Predicted High-Throughput Free-Energy Phase Diagrams for the Discovery of Solid-State Hydrogen Storage Reactions

The design of multinary solid-state material systems that undergo reversible phase changes via changes in temperature and pressure provides a potential means of safely storing hydrogen. However, fully mapping the stabilities of known or newly targeted compounds relative to competing phases at reaction conditions has previously required many stringent experiments or computationally demanding calculations of each compound’s change in Gibbs energy with respect to temperature, G(T). Here, we have extended the approach of constructing chemical potential phase diagrams based on ΔG f (T) to enable the analysis of phase stability at non-zero temperatures. We first performed density functional theory calculations to compute the formation enthalpies of binary, ternary, and quaternary compounds within several compositional spaces of current interest for solid-state hydrogen storage. Temperature effects on solid compound stability were then accounted for using our recently introduced machine learned descriptor for the temperature-dependent contribution G δ (T) to the Gibbs energy G(T). From these Gibbs energies, we evaluated each compound’s stability relative to competing compounds over a wide range of conditions and show using chemical potential and composition phase diagrams that the predicted stable phases and H2 release reactions are consistent with experimental observations. This demonstrates that our approach rapidly computes the thermochemistry of hydrogen release reactions for compounds at sufficiently high accuracy relative to experiment to provide a powerful framework for analyzing hydrogen storage materials. This framework based on G(T) enables the accelerated discovery of active materials for a variety of technologies that rely on solid-state reactions involving these materials.

08 HYDROGEN↗

Machine Learning and IAST-Aided High-Throughput Screening of Cationic and Silica Zeolites for Alkane Capture, Storage, and Separations

We present an approach for quantitatively predicting the temperature-dependent single-component adsorption behavior of linear alkanes in silica and Na-exchanged cationic zeolites using machine learning (ML) models trained from extensive molecular simulations based on force fields with coupled cluster accuracy. A high-performing classification model was developed to distinguish between instances with negligible and non-negligible adsorption. Subsequently, two ML models were trained to predict the single-component adsorption loading and the heat of adsorption at any pressure at 300 K for any zeolite topology and silicon-to-aluminum ratio. The ML models were trained on International Zeolite Association (IZA) zeolites, and their transferability to hypothetical zeolites was successfully validated. We then expand the power of these predictions to adsorbed mixtures at arbitrary temperatures by integrating them with the Clausius–Clapeyron equation and ideal adsorbed solution theory (IAST). This approach was validated and then applied to a temperature swing adsorption separation process to demonstrate its practical utility. We demonstrate how predictions from this ML-enabled approach can allow the selection of high-performing materials that are then validated using detailed molecular simulations based on quantitatively accurate force fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Sim-Situ: A Framework for the Faithful Simulation of in situ Processing

The amount of data generated by numerical simulations in various scientific domains led to a fundamental redesign of how the analysis and visualization of simulation outputs are performed. The throughput and capacity of storage subsystems have not evolved as fast as the computing power in extreme-scale supercomputers, making the classical post-hoc approach highly inefficient. In situ processing has then emerged as a solution in which simulation and data analysis/visualization are intertwined for better performance and greater interactivity.Determining the best allocation, i.e., how many resources to allocate to simulation and analysis respectively, mapping, i.e., where and at which frequency to run the analysis/visualization, and data transfer mode is a complex task whose performance assessment is crucial to the efficient execution of in situ processing. However, such a performance evaluation of different strategies usually relies either on directly running them on the targeted execution environments, which can rapidly become extremely time- and resource-consuming, or on resorting to simplified models of the components of an in situ application, which can lack of realism. In both cases, the validity of the performance evaluation is limited.In this paper, we present Sim-Situ, a simulation-based framework for the faithful performance evaluation of in situ processing strategies. We designed Sim-Situ to reflect the typical features of in situ processing systems. Thanks to its modular design, Sim-situ has the necessary flexibility to easily and faithfully evaluate the behavior and performance of various allocation, mapping, and data transfer strategies. We illustrate the simulation capabilities of Sim-Situ on a Molecular Dynamics use case. We study the impact of different strategies on performance and show how users can leverage Sim-Situ to determine interesting tradeoffs when adding analysis/visualization components to their application.

Honoré, Valentin↗

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache↗

Design principles for the ultimate gas deliverable capacity material: nonporous to porous deformations without volume change

Understanding the fundamental limits of gas deliverable capacity in porous materials is of critical importance as it informs whether technical targets (e.g., for on-board vehicular storage) are feasible. High-throughput screening studies of rigid materials, for example, have shown they are not able to achieve the original ARPA-E methane storage targets, yet an interesting question remains: what is the upper limit of deliverable capacity in flexible materials? In this work we develop a statistical adsorption model that specifically probes the limit of deliverable capacity in intrinsically flexible materials. The resulting adsorption thermodynamics indicate that a perfectly designed, intrinsically flexible nanoporous material could achieve higher methane deliverable capacity than the best benchmark systems known to date with little to no total volume change. Density functional theory and grand canonical Monte Carlo simulations identify a known metal–organic framework (MOF) that validates key features of the model. Therefore, this work (1) motivates a continued, extensive effort to rationally design a porous material analogous to the adsorption model and (2) calls for continued discovery of additional high deliverable capacity materials that remain hidden from rigid structure screening studies due to nominal non-porosity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Virtual Log-Structured Storage for High-Performance Streaming

Over the past decade, given the higher number of data sources (e.g., Cloud applications, Internet of things) and critical business demands, Big Data transitioned from batch-oriented to real-time analytics. Stream storage systems, such as Apache Kafka, are well known for their increasing role in real-time Big Data analytics. For scalable stream data ingestion and processing, they logically split a data stream topic into multiple partitions. Stream storage systems keep multiple data stream copies to protect against data loss while implementing a stream partition as a replicated log. This architectural choice enables simplified development while trading cluster size with performance and the number of streams optimally managed. This paper introduces a shared virtual log-structured storage approach for improving the cluster throughput when multiple producers and consumers write and consume in parallel data streams. Stream partitions are associated with shared replicated virtual logs transparently to the user, effectively separating the implementation of stream partitioning (and data ordering) from data replication (and durability). We implement the virtual log technique in the KerA stream storage system. When comparing with Apache Kafka, KerA improves the cluster ingestion throughput by up to 4x when multiple producers write over hundreds of data streams.

consistent stream ordering↗

Integration of RNTuple in ATLAS Athena

After using ROOT’s TTree I/O subsystem for over two decades and storing more than an exabyte of compressed High Energy Physics (HEP) data, advances in technology have motivated a complete redesign, RNTuple, which breaks backward-compatibility to take better advantage of these storage options. The RNTuple I/O subsystem has been designed to address performance bottlenecks and other shortcomings of TTree. Specifically, RNTuple comes with an updated, more compact binary data format that can be stored both in ROOT files and natively in object stores. It is designed for modern storage hardware (e.g. high-throughput low-latency NVMe SSDs), and provides robust and easy to use interfaces. The binary format of RNTuple is scheduled to become production grade in 2024, and recently has become mature enough to start exploring the integration into software used by HEP experiments. In this contribution, we discuss the developments to support the features as required by the ATLAS analysis Event Data Model (EDM) in RNTuple, which will enable its integration into the Athena software framework. With these developments in place, we evaluate the performance of the current most recent versions of RNTuple-based ATLAS data sets and compare this to that of TTree.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Timely Reporting of Heavy Hitters Using External Memory

Given an input stream S of size N, a Φ-heavy hitter is an item that occurs at least ΦN times in S. The problem of finding heavy-hitters is extensively studied in the database literature. In this work, we study a real-time heavy-hitters variant in which an element must be reported shortly after we see its T = Φ N-th occurrence (and hence it becomes a heavy hitter). We call this the Timely Event Detection (TED) Problem. The TED problem models the needs of many real-world monitoring systems, which demand accurate (i.e., no false negatives) and timely reporting of all events from large, high-speed streams with a low reporting threshold (high sensitivity). Like the classic heavy-hitters problem, solving the TED problem without false-positives requires large space (Ω (N) words). Thus in-RAM heavy-hitters algorithms typically sacrifice accuracy (i.e., allow false positives), sensitivity, or timeliness (i.e., use multiple passes). We show how to adapt heavy-hitters algorithms to external memory to solve the TED problem on large high-speed streams while guaranteeing accuracy, sensitivity, and timeliness. Our data structures are limited only by I/O-bandwidth (not latency) and support a tunable tradeoff between reporting delay and I/O overhead. With a small bounded reporting delay, our algorithms incur only a logarithmic I/O overhead. We implement and validate our data structures empirically using the Firehose streaming benchmark. Multi-threaded versions of our structures can scale to process 11M observations per second before becoming CPU bound. In comparison, a naive adaptation of the standard heavy-hitters algorithm to external memory would be limited by the storage device’s random I/O throughput, i.e., ≈100K observations per second.

97 MATHEMATICS AND COMPUTING↗

Life‐Cycle Assessment Considerations for Batteries and Battery Materials

Abstract Rechargeable batteries are necessary for the decarbonization of the energy systems, but life‐cycle environmental impact assessments have not achieved consensus on the environmental impacts of producing these batteries. Nonetheless, life cycle assessment (LCA) is a powerful tool to inform the development of better‐performing batteries with reduced environmental burden. This review explores common practices in lithium‐ion battery LCAs and makes recommendations for how future studies can be more interpretable, representative, and impactful. First, LCAs should focus analyses of resource depletion on long‐term trends toward more energy and resource‐intensive material extraction and processing rather than treating known reserves as a fixed quantity being depleted. Second, future studies should account for extraction and processing operations that deviate from industry best‐practices and may be responsible for an outsized share of sector‐wide impacts, such as artisanal cobalt mining. Third, LCAs should explore at least 2–3 battery manufacturing facility scales to capture size‐ and throughput‐dependent impacts such as dry room conditioning and solvent recovery. Finally, future LCAs must transition away from kg of battery mass as a functional unit and instead make use of kWh of storage capacity and kWh of lifetime energy throughput.

25 ENERGY STORAGE↗

Leveraging Temperature-Dependent (Electro)Chemical Kinetics for High-Throughput Flow Battery Characterization

The library of redox-active organics that are potential candidates for electrochemical energy storage in flow batteries is exceedingly vast, necessitating high-throughput characterization of molecular lifetimes. Demonstrated extremely stable chemistries require accurate yet rapid cell cycling tests, a demand often frustrated by time-denominated capacity fade mechanisms. We have developed a high-throughput setup for elevated temperature cycling of redox flow batteries, providing a new dimension in characterization parameter space to explore. We utilize it to evaluate capacity fade rates of aqueous redox-active organic molecules, as functions of temperature. We demonstrate Arrhenius-like behavior in the temporal capacity fade rates of multiple flow battery electrolytes, permitting extrapolation to lower operating temperatures. Collectively, these results highlight the importance of accelerated decomposition protocols to expedite the screening process of candidate molecules for long lifetime flow batteries.

25 ENERGY STORAGE↗