Search NASA⌕ Search

SEARCH · Search NASA

Results for “data storage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Global Corn Heat Stress: Mean and SD of Degree Days Above 29°C based on NEX-GDDP-CMIP6 Climate Projections

Description This global dataset provides the estimated mean and standard deviation (SD) of corn heat stress (degree days above 29°C) for a set of climate models in NEX-GDDP-CMIP6 at 0.25-degree resolution. The NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6). The current dataset includes: Long-Term Average Degree Days Above 29°C- Historical Long-Term Average Degree Days Above 29°C- SSP245 Long-Term Standard Deviation of Degree Days Above 29°C- Historical Long-Term Standard Deviation of Degree Days Above 29°C- SSP245 The mean and SD are calculated over 1985-2014 for the historical period and over 2035-2064 for future projections. A full description of methods, including growing season, daily temperature distribution, and statistical coefficients, can be found in Haqiqi (2024). The source climate data are obtained from https://ds.nccs.nasa.gov/thredds2/catalog/catalog.html and are described in Thrasher et al (2022). The codes used to create this dataset are available at https://github.com/ihaqiqi/dd29c_nex_cmip6. Acknowledgments This work was supported by the US Department of Energy, Office of Science, Biological and Environmental Research Program, Earth and Environmental Systems Modeling, MultiSector Dynamics under Cooperative Agreement DE-SC0022141. The data processing, computation, and storage were completed on Purdue Anvil supercomputer and cyberinfrastructure supported by the National Science Foundation HDR award # 2118329: "NSF Institute for Geospatial Understanding through an Integrative Discovery Environment (I-GUIDE)". References Haqiqi. I. (2024). Trade can buffer climate-induced risks and volatilities in crop supply. Environmental Research: Food Systems. https://doi.org/10.1088/2976-601X/ad7d12 Thrasher, B., Wang, W., Michaelis, A., Melton, F., Lee, T., & Nemani, R. (2022). NASA global daily downscaled projections, CMIP6. Scientific Data, 9(1), 262. https://doi.org/10.1038/s41597-022-01393-4

Climate Change↗

Best Practices for Nuclear Experiment Data Preservation at Idaho National Laboratory: A Guide for Researchers and Reactor Operators

Preserving experimental data is essential for supporting advancements in nuclear science and ensuring the longevity of Idaho National Laboratory's contributions to reactor technology and safety. This report provides a comprehensive guide to best practices for experimental data management and preservation, focusing on standardized data formats, redundancy in storage, metadata documentation, and alignment with international standards. By following these recommendations, experimentalists and reactor operators can enhance the accessibility, reproducibility, and utility of critical datasets for regulatory review, validation computational methods, and future research.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine learning-guided design of direct methanol fuel cells with a platinum group metal-free cathode

Direct methanol fuel cells (DMFCs) offer a promising solution for clean electricity generation, particularly in small electronics and remote auxiliary power units. However, optimizing their efficiency and performance is challenging due to the complex interactions between various factors. Here, we present a novel approach that integrates experiments with machine learning to model and predict the performance of these fuel cells using atomically dispersed platinum group metal (PGM)-free catalysts at the cathode. Further, our machine learning models, trained on diverse input parameters, allow for the comprehensive optimization of DMFC performance prior to fabrication and testing. Through extensive experimental validation, we demonstrate that this data-driven approach accurately predicts key performance metrics, such as maximum power output and polarization curves. By combining our models with interpretable game-theory methods, we provide deep insights into the factors governing fuel cell performance, ultimately paving the way for the design of scalable and efficient DMFC technologies.

25 ENERGY STORAGE↗

Carbon Storage Site Mapping Inquiry Tool (MapIT)

To date, 48 projects, consisting of 139 wells, are currently under review with the Environmental Protection Agency’s (EPA) Underground Injection Control (UIC) Program for Class VI – wells used for geologic sequestration of carbon dioxide. The number of applications submitted is expected to increase in coming years with the increase of the 45Q tax credit available to projects that initiate construction prior to 2033. The amount of data collected to submit a Class VI permit is vast, and often disparate, coming from state, federal, and commercial entities, as well as field-specific data collected within an area of interest. When preparing for site selection and permitting, the initial aggregation of relevant public data can be time intensive. The Carbon Storage Site Mapping Inquiry tool (MapIT) was created to support and accelerate the discovery and accessibility of open-source data and information available across the USA. Data was aggregated and organized based on data types described within the EPA UIC Class VI permit documentation. The online tool enables users to explore hundreds of geospatial data layers and connect to additional external resources, leveraging API and REST services where possible to ensure updates to data in real time. MapIT enables users to explore state and federal data related to geologic, geophysical, structural, hydrologic, and contextual information. In addition to displaying spatial data and linking to external resources, MapIT leverages custom widgets to ensure that internal data and external data are discoverable and accessible. The widgets connect users to resources such as the USGS publications and the USGS Earthquake Catalog based on a user-defined location. This talk will describe data aggregation workflows, data types, data preparation, and tool development for MapIT. The Carbon Storage Site Mapping Inquiry Tool and underlying database are valuable, intuitive resources that empower government, academic, commercial and industry stakeholders to explore, analyze, and acquire carbon storage related data.

Morkner, Paige↗

Carbon Transport and Storage Planning and Viability Support Tools

The EDX disCO2ver Carbon Transport and Storage Planning and Viability Support Tools are made up of the Carbon Storage Planning Inquiry Tool (CS PlanIT, Justman et al. 2024) and the Carbon Storage Technical Viability Approach Support Tool (CS TVA). Together, these tools support data access to support understanding data availability to support planning efforts for carbon transport and storage. The Carbon Storage Planning Inquiry Tool (CS PlanIT) is an online web mapping application designed to help users explore, query, and evaluate multiple data layers to support and accelerate carbon storage resource and feasibility assessments and planning efforts. CS PlanIT currently contains a range of datasets associated with geologic, technical, and infrastructure factors. The data sets can be filtered geographically for an area of interest to update statistics and charts within the dashboard. The dashboard is divided into different sections called widgets, relating to different steps in the carbon storage planning process. The resources in this submission include a link to PlanIT, as well as a data catalog and link to user documentation. The original citation for the CS PlanIT tool, which has now been integrated into the toolset here, was: - Devin Justman, Scott Pantaleone, Maneesh Sharma, Lucy Romeo, Paige Morkner, CS PlanIT (Carbon Storage Planning Inquiry Tool) , 6/28/2024, https://edx.netl.doe.gov/dataset/cs-planit-carbon-storage-planning-inquiry-tool, DOI: 10.18141/2377953 The Carbon Storage Technical Viability Approach Support (CS TVA) Tool displays spatial data availability for the many components of Geologic Carbon Storage (GCS) technical viability assessment (Creason al 2025). Identifying sites suitable for GCS requires evaluating the intersection of myriad factors, including reservoir conditions, subsurface and surface hazards, infrastructure requirements, and energy community metrics. The technical viability of a site can only be confirmed for instances where all these factors have data available, and where those data support viability. Additional Resources related to the Technical Viability Assessment Tool: - Julia Mulhern, Casey White, Araceli Lara, Neyda Cordero Rodriguez, Zachary Jackson, Jacob Shay, Gabriel Creason, MacKenzie Mark-Moser, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Database, 3/26/2025, https://edx.netl.doe.gov/dataset/edx4ccs-carbon-storage-technical-viability-approach-database , DOI:10.18141/1984655 - Julia Mulhern, MacKenzie Mark-Moser, Gabriel Creason, Casey White, Araceli Lara, Neyda Cordero Rodriguez, Zach Jackson, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Matrix, 3/27/2025, https://edx.netl.doe.gov/dataset/carbon-storage-technical-viability-approach-cs-tva-matrix , DOI: 10.18141/2539979 - Gabriel Creason, Zach Jackson, Neyda Cordero Rodriguez, Julia Mulhern, Casey White, Araceli Lara, MacKenzie Mark-Moser, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Data Availability Results Database, 3/27/2025, https://edx.netl.doe.gov/dataset/carbon-storage-technical-viability-approach-cs-tva-data-availability-results-database, DOI:10.18141/2538557

Carbon storage↗

Overview of the distributed image processing infrastructure to produce the Legacy Survey of Space and Time

The Vera C. Rubin Observatory is preparing to execute the most ambitious astronomical survey ever attempted, the Legacy Survey of Space and Time (LSST). Currently the final phase of construction is under way in the Chilean Andes, with the Observatory’s ten-year science mission scheduled to begin in 2025. Rubin’s 8.4-meter telescope will nightly scan the southern hemisphere collecting imagery in the wavelength range 320–1050 nm covering the entire observable sky every 4 nights using a 3.2 gigapixel camera, the largest imaging device ever built for astronomy. Automated detection and classification of celestial objects will be performed by sophisticated algorithms on high-resolution images to progressively produce an astronomical catalog eventually composed of 20 billion galaxies and 17 billion stars and their associated physical properties. In this article we present an overview of the system currently being constructed to perform data distribution as well as the annual campaigns which reprocess the entire image dataset collected since the beginning of the survey. These processing campaigns will utilize computing and storage resources provided by three Rubin data facilities (one in the US and two in Europe). Each year a Data Release will be produced and disseminated to science collaborations for use in studies comprising four main science pillars: probing dark matter and dark energy, taking inventory of solar system objects, exploring the transient optical sky and mapping the Milky Way. Also presented is the method by which we leverage some of the common tools and best practices used for management of large-scale distributed data processing projects in the high energy physics and astronomy communities. We also demonstrate how these tools and practices are utilized within the Rubin project in order to overcome the specific challenges faced by the Observatory.

79 ASTRONOMY AND ASTROPHYSICS↗

LTE Electrolyzer Data Collection

The goal for NREL is to collect, develop and publish performance metrics relative to low temperature electrolyzer installations. This will be done through the development of: Secure storage solution to house the collection of data from multiple projects Standardization of data to be collected and analyzed. This will be done using data templates developed with the help of partners involved with electrolyzer installations. Analysis that produces metrics of interest for all stakeholders Aggregation of results from multiple projects to view industry progress as a whole Publication of aggregated results in the form of composite data products (CDPs) Collaboration with Idaho National Lab and their work with high temperature electrolyzer installations will enable efficient use of storage and analysis tools.

data↗

Orbital Engineering Band Degeneracy in a Dual-Square Carbon-Oxide Framework

Electron band degeneracies in momentum space give rise to exotic quantum phenomena that have sparked intense interest in condensed matter physics and materials science. Nodal-lines─isolines in k-space formed by the incidental touching of two bands that share the same energy but belong to discrete eigenstates─arise in the presence of symmetries that preclude effective hybridization. Despite recent advances in the design, bottom-up assembly, and engineering of exotic electronic states in graphene nanomaterials, the extension of this approach to access synthetic two-dimensional (2D) quantum materials derived from metal- or covalent-organic frameworks (COFs) has lagged behind. Here we present a molecular orbital engineering approach for designing and fabricating an edge-centered dual square lattice within a π-conjugated 2D-tetraoxa[8]circulene (2D-TOC) COF. First-principles calculations and scanning tunnelling spectroscopy reveal the emergence of Frontier states at the center of a 3 × 3 lattice that give rise to Dirac nodal-lines in 2D-TOC. Our findings not only provide a general guide for the design of conjugated COFs with custom tailored electronic properties from molecular fragments but enable the exploration of emergent topological phenomena in synthetic 2D materials with potential application for high-speed, low-power data processing, transmission, and storage.

Dirac nodal line semimetal↗

Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures

Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.

kilic, Ozgur Ozan [Brookhaven National Laboratory ↗

Opportunities and challenges to study solar neutrinos with a Q-Pix pixel readout

The study of solar neutrinos presents significant opportunities in astrophysics, nuclear physics, and particle physics. However, the low-energy nature of these neutrinos introduces considerable challenges to isolate them from background events, requiring detectors with low-energy threshold, high spatial and energy resolutions, and low data rate. We present the study of solar neutrinos with a kiloton-scale liquid argon detector located underground, instrumented with a pixel readout using the Q-Pix technology. We explore the potential of using volume fiducialization, directional topological information, light signal coincidence, and pulse-shape discrimination to enhance solar neutrino sensitivity. We find that discriminating neutrino signals below 5 MeV is very difficult. However, we show that these methods are useful for the detection of solar neutrinos when external backgrounds are sufficiently understood and when the detector is built using low-background techniques. When building a workable background model for this study, we identify 𝛾 background from the cavern walls and from capture of 𝛼 particles in radon decay chains as both critical to solar neutrino sensitivity and significantly underconstrained by existing measurements. Finally, we highlight that the main advantage of the use of Q-Pix for solar neutrino studies lies in its ability to enable the continuous readout of all low-energy events with minimal data rates and manageable storage for further off-line analyses.

multi-purpose particle detectors↗

An Efficient Checkpointing System for Large Machine Learning Model Training

As machine learning models increase in size and complexity rapidly, the cost of checkpointing in ML training became a bottleneck in storage and performance (time). For example, the latest GPT-4 model has massive parameters at the scale of 1.76 trillion. It is highly time and storage consuming to frequently writes the model to checkpoints with more than 1 trillion floating point values to storage. This work aims to understand and attempt to mitigate this problem. First, we characterize the checkpointing interface in a collection of representative large machine learning/language models with respect to storage consumption and performance overhead. Second, we propose the two optimizations: i) A periodic cleaning strategy that periodically cleans up outdated checkpoints to reduce the storage burden; ii) A data staging optimization that coordinates checkpoints between local and shared file systems for performance improvement.

machine learning, artificial intelligence↗

Application of Machine Learning and Data Augmentation Algorithms in the Discovery of Metal Hydrides for Hydrogen Storage

The development of efficient and sustainable hydrogen storage materials is a key challenge for realizing hydrogen as a clean and flexible energy carrier. Among various options, metal hydrides offer high volumetric storage density and operational safety, yet their application is limited by thermodynamic, kinetic, and compositional constraints. In this work, we investigate the potential of machine learning (ML) to predict key thermodynamic properties—equilibrium plateau pressure, enthalpy, and entropy of hydride formation—based solely on alloy composition using Magpie-generated descriptors. We significantly expand an existing experimental dataset from ~400 to 806 entries and assess the impact of dataset size and data augmentation, using the PADRE algorithm, on model performance. Models including Support Vector Machines and Gradient Boosted Random Forests were trained and optimized via grid search and cross-validation. Results show a marked improvement in predictive accuracy with increased dataset size, while data augmentation benefits are limited to smaller datasets and do not improve accuracy in underrepresented pressure regimes. Furthermore, clustering and cross-validation analyses highlight the limited generalizability of models across different material classes, though high accuracy is achieved when training and testing within a single hydride family (e.g., AB2). The study demonstrates the viability and limitations of ML for accelerating hydride discovery, emphasizing the importance of dataset diversity and representation for robust property prediction.

augmentation↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

Building partnerships for development of sustainable energy systems with atmospheric measurements

Atmospheric dynamics often play a critical role in the sustainability and reliability of diverse forms of energy production. This is especially true for the growing number of renewable energy deployments that harness aspects of the environment for power production. While the University of Memphis has a strong research background in energy systems, we have little experience working with the Earth and Environmental Systems Science Division (EESSD) and their associated User Facilities. Of particular interest to us is the Atmospheric Science Research and the Atmospheric Radiation Measurement (ARM) user facility to address surface-boundary layer interactions and physical phenomena. One of the major challenges for understanding and developing energy systems and management platforms is accurate modeling/forecasting of atmospheric conditions across disparate spatial and temporal scales. These conditions are often required to understand the lowest levels of the atmospheric boundary layer, but are also important to understand higher atmospheric conditions where aerosols affect cloud development. The objective of this work was to develop partnerships with national laboratories for collaboration on environmental science and its intersection with sustainable energy systems, as well as to leverage the ARM user facility data repositories to enhance our research capabilities in energy systems and their inter-dependence on environmental systems for future engagement with EESSD. Specifically, we accomplished these objectives by (1) developing collaborations with Oakridge National Laboratory ARM Data Science and Integration Group which resulted in student internships, (2) employed ARM data to develope modeling of the atmospheric boundary layer optical turbulence, and (3) optimally-sized large-scale renewable energy systems and their associated energy storage systems with ARM repository data.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Lithium-Ion Battery Design for Grid-Scale Energy Storage App

A software that delivers parameters from energy storage system (ESS) to container, rack, module and single cell design, as well as data analysis on arbitrage energy and frequency regulation of ESS in different regions, has been developed. The Lithium-ion Battery Design for Grid-scale Energy Storage App V1.0 has the capability to output the system, module and cell design with the energy, power, capacity, group method, cost of single cell, and single cell test protocol which break down from input energy storage system data in different regions. The default chemistry of the battery is LiFePO4 and graphite. The energy density of the graphite/LiFePO 4 pouch cell ranges from 100 Wh/kg to 200 Wh/Kg in the software. Graphite/LiFePO 4 pouch cell (up to 1Ah in lab) manufacturing line is also built and can be used to evaluate the test protocol, moreover, for electrolyte evaluation in other ESMI seedling projects. The software enables rapid prototyping to accelerate energy storage research, development, and manufacturing.

Liu, Dianying [Pacific Northwest National Laborato↗

A General Framework for Error-controlled Unstructured Scientific Data Compression

Data compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the −6 to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets.

Gong, Qian↗