Search NASASearch

SEARCH · Search NASA

Results for “data lake”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluate data lake design for the accelerator control system

Increasing precision in automation for modern particle accelerators not only creates a requirement to gather data from all devices but also demands scalable and high-performance data infrastructure with the capability of handling vast incoming device data. A well architected data lake is suitable for such a system which integrates real-time data acquisition, transient data caching, and long-term storage. This paper evaluates data lake architecture for an Accelerator Control System (ACS), focusing on two critical components of a data lake, data cache and long-term storage.

Jaikar, Amol [Fermilab]

Evaluation of Data Lake Design for the Accelerator Control System

In this modern world, the user expects faster processing and real-time response to operate accelerator control devices. The existing framework with its infrastructure does not have the ability to satisfy these future needs. Therefore, modernization is required to reach industry standards and develop a modular framework that can provide flexibility and dynamic scalability. The data lake architecture, comprising three layers, ingestion, processing, and data consumption, provides flexibility and scalability to meet current and future demands for the accelerator control system.

Jaikar, Amol [Fermilab]

Common Column Identification for Table Similarity Detection in Electrified Transportation Data Lakes

Electrified transportation often requires researchers and operators to interact with datasets from a wide range of sources and disciplines, such as transportation, power systems, public health, policies, and regulations. These datasets vary in quality and format, making it difficult to understand, preprocess, and identify key columns representing real-world entities or values for indexing and joining, which can negatively impact downstream analysis and operation. Existing solutions are limited, requiring extensive manual customization or data expertise to utilize. In this article, we propose a multi-layered approach to automatically identify key columns to expedite preprocessing and aid in analysis of electrified transportation data. Our method leverages a dynamic ontology to identify common fields and an information theory-based strategy for edge cases that are difficult to generalize. Evaluations on a number of datasets from data.gov and kaggle.com show improved performance of our methods over several baseline techniques, and our ablation analyses illustrate the efficacy of individual components of our method. Our case studies also demonstrate that our methods have the potential to improve analysis of electrified transportation data and aid in automatic integration of such datasets.

33 ADVANCED PROPULSION SYSTEMS

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility

The Zooplankton International Geospatial dataset: A global repository of spatiotemporal freshwater zooplankton community composition data from lakes and reservoirs to support ecological research

Zooplankton transfer substantial energy in aquatic food webs and are used as indicators of environmental change. Syntheses of zooplankton community dynamics globally require datasets that span a wide range of environmental gradients; however, these datasets are limited due to methodological differences across programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we created the Zooplankton International Geospatial (ZIG) dataset, which includes original zooplankton, water physical and chemical variables, and lake morphometric data from 311 inland lakes and reservoirs. ZIG includes waterbodies ranging in size from 0.005 to 82,100 km2 and spanning broad latitudinal (−47.26 to 64.90) and longitudinal ranges (−165.04 to 176.53). Temporal coverage for individual waterbodies ranges between 1 and 60 yr with sampling frequency ranging from annually to weekly. With its extensive coverage and content, we consider ZIG to be a cornerstone for future investigations of global scale lake biodiversity change.

Figary, Stephanie [Cornell University, Ithaca, NY]

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility

Open Energy Data Initiative (OEDI) FY22-24 (Final Technical Report)

Final technical report for the Open Energy Data Initiative (OEDI) project covering fiscal years FY22 through FY24. The DOE Open Energy Data Initiative (OEDI) is a partnership between the National Renewable Energy Laboratory (NREL), the U.S. Department of Energy (DOE), and major cloud providers including Amazon, Microsoft, and Google to provide universal access to big data in the cloud. At the heart of OEDI is a centralized repository of high-value energy research datasets aggregated from the U.S. Department of Energy's Program Offices, National Laboratories and other collaborators. It aggregates smaller, domain-specific repositories, allows direct data submissions, and includes support for big data through its energy data lakes. OEDI's data lakes make high-value data universally accessible and help researchers, collaborators and the general public overcome many of the obstacles to accessing and using big data.

29 ENERGY PLANNING, POLICY, AND ECONOMY

2024 Annual Technology Baseline (ATB) Cost and Performance Data for Electricity Generation Technologies

These data provide the 2024 update of the Electricity Annual Technology Baseline (ATB). Starting in 2015 NREL has presented the ATB, consisting of detailed cost and performance data, both current and projected, for electricity generation and storage technologies. The ATB products now include data (Excel workbook, Tableau workbooks, and structured summary csv files), as well as documentation and user engagement via a website, presentation, and webinar. Starting in 2021, the data are cloud optimized and provided in the OEDI data lake. The data for 2015 - 2020 are can be found on the NREL Data Search Page. The website documentation can be found on the ATB Website.

Array

Soil and groundwater environmental sensor data, Wax Lake Delta, Louisiana, March 2023 - March 2024

This study evaluates how environmental parameters that integrate biogeochemical processes vary with water table fluctuations in the freshwater Wax Lake Delta (WLD) in Louisiana, U.S.A. This data package contains seven *.csv files and one Excel file that compiles all the data from the individual .csv files. This dataset reports high frequency (15-min) observations of water level, soil redox potential, specific conductance, and pH made for one year along elevation transects located on the older, proximal (OT) and younger, distal (YT) ends of a deltaic island. Water depth relative to the ground surface (cm; HOBO U20L-04; error ± 0.4 cm), water pH and temperature (HOBO MX2501), and specific conductance and temperature (HOBO U24-001) sensors were installed in March 2023. Water depth was corrected for barometric pressure recorded by a separate logger secured to a platform above the highest water level. Soil redox probes (SWAP ORP-40-4-B) were also installed in March 2023. Each probe had four Pt sensors (2 mm width) placed at 10 cm, 20 cm, 30 cm, and 40 cm below the ground surface. Redox data were referenced to an external Ag0/AgCl (3M KCl) reference probe placed in saturated ground and recorded on CR1000X dataloggers (Campbell Scientific) powered by solar panels. A second reference probe was positioned near the primary reference probe for backup and data correction. The tops of the soil redox probes and soil moisture probes were flush with the soil surface so that sensors are reported at their indicated depths below ground surface. Here, we report data collected between 15 March 2023 to 15 March 2024 for all sensors, with some differences due to exact dates of sensor placement or data gaps associated with sensor malfunction. For example, water depth at OT4 was not recorded between March to November 2023. Data flags indicate whether a value is valid (1) or was excluded from data analysis in the associated manuscript (-1).

EARTH SCIENCE > LAND SURFACE > SOILS

ComStock Measure Scenario Documentation: Chiller Replacement

Building on a 3-year effort to calibrate and validate the U.S. Department of Energy's ResStock (TM) and ComStock (TM) models, this work produces national datasets that empower analysts working for federal, state, utility, city, and manufacturer stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual energy consumption (at a subhourly resolution) of the commercial building stock across the United States. The baseline model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology and results of the baseline model are discussed in the final technical report of the End-Use Load Profiles project. The goal of this work is to develop energy efficiency and demand flexibility end-use load shapes that cover high-impact, market-ready (or nearly market-ready) measures. "Measures" refers to various "what-if" scenarios that can be applied to buildings. An end-use savings shape is the difference in energy consumption between a baseline building (or collection of buildings) and a building with an energy efficiency or demand flexibility measure applied. It results in a time-series profile broken down by end use and fuel (electricity or on-site gas, propane, or fuel oil use) at each time step, as well as annual aggregations. This report describes the modeling methodology for a single end-use savings shape measure - chiller replacement - and briefly introduces key results. The full public dataset can be accessed on the ComStock (TM) data lake or via the Data Viewer at comstock.nrel.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

ComStock Measure Documentation: High-Efficiency Rooftop Unit

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - high-efficiency rooftop unit (RTU) - and briefly introduces key results. The full public data set can be accessed on the Comstock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

ComStock Measure Documentation: Variable-Speed Pumps

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - variable speed pumps - and briefly introduces key results. The full public data set can be accessed on the ComStock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

ComStock Measure Documentation: Fan Static Pressure Reset for Multizone Variable Air Volume Systems

This report assesses the potential for nationwide adoption of a duct static pressure reset in MZ VAV systems in appropriate applications. Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy’s ResStock™ and ComStock™ models, this work produces national datasets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual sub-hourly energy consumption of the commercial building stock across the United States. The “baseline” model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass-adoption impact on the baseline building stock. “Measures” refers to various “what-if” scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public datasets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario—Fan Static Pressure Reset for Multizone Variable Air Volume (VAV) Systems—and briefly introduces key results. The full public dataset can be accessed on the ComStock data lake or via the Data Viewer at comstock.nrel.gov. The public dataset enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

ComStock Measure Documentation: Thermostat Setbacks During Unoccupied Periods

This report assesses the potential for nationwide adoption of thermostat setbacks in appropriate applications. Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy’s ResStock™ and ComStock™ models, this work produces national datasets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual sub-hourly energy consumption of the commercial building stock across the United States. The “baseline” model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass-adoption impact on the baseline building stock. “Measures” refers to various “what-if” scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public datasets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario— Thermostat Setbacks During Unoccupied Periods—and briefly introduces key results. The full public dataset can be accessed on the ComStock data lake or via the Data Viewer at comstock.nrel.gov. The public dataset enables users to create custom aggregations of results for their use cases (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

ResStock Measure Documentation: Cold Climate Air-Source Heat Pump

This report is part of series describing a variety of different ResStock™ measures. "Measures" refers to energy efficiency retrofits that can be applied to buildings during modeling. This documentation covers the "Cold Climate Air-Source Heat Pump" measure upgrade methodology and briefly discusses key results. All results can be accessed on the ResStock Open Energy Data Initiative "End-Use Load Profiles for the U.S. Building Stock" data lake and on the data viewer at resstock.nlr.gov.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

ResStock Measure Documentation: Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER) With Light Envelope Improvements

This report is part of series describing a variety of different ResStock(TM) measures. "Measures" refers to energy efficiency retrofits that can be applied to buildings during modeling. This documentation covers the "Residential Two-Stage Geothermal Heat Pump (4.0 COP, 20.5 EER) With Light Envelope Improvements" measure upgrade methodology and briefly discusses key results. All results can be accessed on the ResStock Open Energy Data Initiative "End-Use Load Profiles for the U.S. Building Stock" data lake and on the data viewer at resstock.nlr.gov.

15 GEOTHERMAL ENERGY

ResStock Measure Documentation: Residential Variable-Speed Geothermal Heat Pump (4.4 COP, 30.9 EER) With Light Envelope Improvements

This report is part of series describing a variety of different ResStock™ measures. "Measures" refers to energy efficiency retrofits that can be applied to buildings during modeling. This documentation covers the "Residential Variable-Speed Geothermal Heat Pump (4.4 COP, 30.9 EER) with Light Envelope Improvements" measure upgrade methodology and briefly discusses key results. All results can be accessed on the ResStock Open Energy Data Initiative "End-Use Load Profiles for the U.S. Building Stock" data lake and on the data viewer at resstock.nlr.gov.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI