Search NASA⌕ Search

SEARCH · Search NASA

Results for “Building Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Building Datasets and Training Methods for ML Based Magnet Quench Detection

Detecting quenches in superconducting (SC) magnets during training is a challenging process that involves capturing physical events that occur at different frequencies and appear as various signal features. These events may be correlated across instrumentation type, thermal cycle, and ramp. These events together build a more complete picture of continuous processes occurring in the magnet, and may allow us to flag potential precursors for quench detection. We present our work on building an automatic machine learning (ML) based quench detection system. We build upon our existing work on unsupervised auto-encoders for acoustic sensors and quench antenna (QA) by first establishing a supervised ML training pipeline. We show the results of an event tagging, analysis, and simulation framework on our QA and acoustic data which are used concurrently to build a training dataset for a supervised implementation. We then show how this supervised training can be used as a prior in a semi-supervised framework and compare this to the unsupervised neural network auto-encoder performance.This allows us to have a more concrete understanding of the performance of our algorithms relative to physical events occurring in the magnet, and also provides a baseline software tool to generically evaluate our quench prediction autoencoders under completely unsupervised, supervised, and semi-supervised training conditions.

Khan, Maira [Fermilab]↗

A labeled dataset for building HVAC systems operating in faulted and fault-free states

Abstract Open data is fueling innovation across many fields. In the domain of building science, datasets that can be used to inform the development of operational applications - for example new control algorithms and performance analysis methods - are extremely difficult to come by. This article summarizes the development and content of the largest known public dataset of building system operations in faulted and fault free states. It covers the most common HVAC systems and configurations in commercial buildings, across a range of climates, fault types, and fault severities. The time series points that are contained in the dataset include measurements that are commonly encountered in existing buildings as well as some that are less typical. Simulation tools, experimental test facilities, and in-situ field operation were used to generate the data. To inform more data-hungry algorithms, most of the simulated data cover a year of operation for each fault-severity combination. The data set is a significant expansion of that first published by the lead authors in 2020.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Building-Level Comparison of Microsoft and Google Open Building Footprints Datasets

Large-scale datasets of building footprints are a crucial source of information for a variety of efforts. In 2023, the general public benefits from open access to multiple sources of building footprints at the country scale or larger, such as those produced by Microsoft and Google. However, none of the available datasets have attained complete global coverage, and researchers and analysts may need to combine multiple sources to assemble a complete set of building footprints for their area of interest or choose between overlapping sources, requiring an understanding of the differences between different building sources. This paper presents a method to closely examine the quality of different building footprint sources by matching corresponding buildings across datasets, using building footprints in Ethiopia published by Microsoft and Google as an example set.

Gonzales, Jack↗

Data-driven key performance indicators and datasets for building energy flexibility: A review and perspectives

Energy flexibility, through short-term demand-side management (DSM) and energy storage technologies, is now seen as a major key to balancing the fluctuating supply in different energy grids with the energy demand of buildings. This is especially important when considering the intermittent nature of ever-growing renewable energy production, as well as the increasing dynamics of electricity demand in buildings. This paper provides a holistic review of (1) data-driven energy flexibility key performance indicators (KPIs) for buildings in the operational phase and (2) open datasets that can be used for testing energy flexibility KPIs. The review identifies a total of 48 data-driven energy flexibility KPIs from 87 recent and relevant publications. These KPIs were categorized and analyzed according to their type, complexity, scope, key stakeholders, data requirement, baseline requirement, resolution, and popularity. Moreover, 330 building datasets were collected and evaluated. Of those, 16 were deemed adequate to feature building performing demand response or building-to-grid (B2G) services. The DSM strategy, building scope, grid type, control strategy, needed data features, and usability of these selected 16 datasets were analyzed. This review reveals future opportunities to address limitations in the existing literature: (1) developing new data-driven methodologies to specifically evaluate different energy flexibility strategies and B2G services of existing buildings; (2) developing baseline-free KPIs that could be calculated from easily accessible building sensors and meter data; (3) devoting non-engineering efforts to promote building energy flexibility, standardizing data-driven energy flexibility quantification and verification processes; and (4) curating and analyzing datasets with proper description for energy flexibility assessm.

24 POWER TRANSMISSION AND DISTRIBUTION↗

2024 Buildings Technology Baseline: Dataset Documentation

The Buildings Technology Baseline is a curated and regularly updated dataset of current and projected performance, retail, and installed price data for all major building energy technologies needed to enable cost/benefit analyses. Building technology analyses require an up-to-date understanding of installation costs and cost-effectiveness of key building energy efficiency technologies. The dataset was assembled by Guidehouse during fiscal year 2024. Data was gathered from the 2024 National Residential Efficiency Measures Database (NREMDB), the 2023 Energy Information Administration Updated Buildings Sector Appliance and Equipment Costs and Efficiencies ("EIA Building Data Report"), DOE Lighting Market Model, the 2023 RSMeans database, and the 2020 Grid-Interactive Efficient Building Technology Cost, Performance, and Lifetime Characteristics ("GEB Data Report"), Lawrence Berkeley National Laboratory, various literature, as well as new data from online retailers, stakeholder interviews, and contractor databases in 2023 and 2024. The dataset has been reviewed by subject matter experts at NREL and DOE. The 2024 dataset release is intended to be a starting point for interested users to provide feedback. This database is not intended to provide specific cost estimates for a specific project. The cost estimates do not include any rebates or tax incentives that may be available for the measures. Rather, it is meant to help determine which measures may be more cost-effective. The National Renewable Energy Laboratory (NREL) makes every effort to ensure accuracy of the data; however, NREL does not assume any legal liability or responsibility for the accuracy or completeness of the information.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Visual Brick model authoring tool for building metadata standardization

In this study, the Brick ontology is a unified semantic metadata standard for building assets and their relationships, serving as a key enabler for effective interoperability and automation of building systems and analytics. However, creating a Brick model, in other words, standard semantic metadata based on the Brick ontology for a building dataset, can be a complex task. This paper presents two case studies of the creation of Brick models for real-world residential and commercial building datasets, highlighting the challenges during the Brick model creation process. Additionally, the paper introduces VizBrick, an interactive authoring tool for creating semantic building metadata. VizBrick facilitates the creation of Brick models by providing an intuitive visual interface and interactive capabilities, such as keyword search, automatic mapping suggestions, and recommendations. The use of VizBrick is shown to significantly reduce the time and effort required during the Brick model creation process.

42 ENGINEERING↗

Mapping use cases and dataset needs for benchmarking buildings data

A perennial challenge in buildings research is the lack of high-quality datasets that can be relied upon for a wide array of tasks, including model calibration and improving energy efficiency and load flexibility. Instrumenting a building for data collection is resource intensive, so it is important to be methodical in the approach and ensure that resulting data are flexible and useful for a broad range of analyses. This study aims to fill the gaps in characterizing potential use cases for buildings datasets and mapping them to dataset needs using a well-defined data infrastructure. Here, we have developed a systematic mapping strategy between buildings dataset needs and use cases to help streamline the processes of efficiently targeting datasets, designing building sensing systems, and determining buildings research use cases. We selected 14 prospective use cases and 11 refined buildings data categories for developing the preliminary dataset-needs-to-use-cases mapping matrix (‘DN-UC mapping matrix’) with generic ‘Tags’—a detailed sub-level of data categories extracted by justifying the needs of an aspect of the datasets to use cases. We present two example applications of the developed mapping matrix to demonstrate use of the mapping matrix and its effectiveness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ComStock: Commercial Building Stock Energy Consumption Dataset

The commercial building sector stock model, or ComStock, is a highly granular, bottom-up model that uses multiple data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States.

building↗

A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data centric approach emphasizes leveraging available data throughout the production process to optimize performance. Integration of extensive data analysis provides the opportunity to improve precision, reduce waste, and enhance the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes the comprehensive description of deposition process, process parameters, in-situ collected welding characteristics, acoustic data, and X-Ray Computed Tomography analysis data for the build. Dataset A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process has arisen under UT-Battelle, LLC’s Prime Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy (DOE) to manage and operate the Oak Ridge National Laboratory. UT-Battelle, LLC will not assert any rights under United States law or under the Prime Contract it has in the dataset against any user of the dataset, including any copyrights or patent rights. UT-Battelle, LLC requests that attribution to the dataset is provided as academically appropriate.

42 ENGINEERING↗

Downscaled Earth System Model Data for Resilient Energy System Planning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. In this presentation, we explore the output characteristics of the dataset and various validation analyses. We also present and discuss plans for the integration of this data into power system planning models using a decision-making under deep uncertainty (DMDU) methodology.

97 MATHEMATICS AND COMPUTING↗

Aircraft Engine Run-to-Failure Dataset Under Real Flight Conditions for Prognostics and Diagnostics

A key enabler of intelligent maintenance systems is the ability to predict the remaining useful lifetime (RUL) of its components, i.e., prognostics. The development of data-driven prognostics models requires datasets with run-to-failure trajectories. However, large representative run-to-failure datasets are often unavailable in real applications because failures are rare in many safety-critical systems. To foster the development of prognostics methods, we develop a new realistic dataset of run-to-failure trajectories for a fleet of aircraft engines under real flight conditions. The dataset was generated with the Commercial Modular Aero-Propulsion System Simulation (CMAPSS) model developed at NASA. The damage propagation modelling used in this dataset builds on the modelling strategy from previous work and incorporates two new levels of fidelity. First, it considers real flight conditions as recorded on board of a commercial jet. Second, it extends the degradation modelling by relating the degradation process to its operation history. This dataset also provides the health respectively fault class. Therefore, besides its applicability to prognostics problems, the dataset can be used for fault diagnostics.

CMAPPS↗

An open retail boundary dataset for South Korea using open data and computer vision technique

Although delineating retail boundaries is important to explore and comprehend the dynamics of the retail sector, it is hard to find studies specifically addressing it in the South Korean context. This study fills this gap by proposing new retail boundaries across South Korea. To achieve this goal, we employed a variety of retailers and building datasets and proposed a unique computer vision-based framework with a deep ensemble voting technique. As a result, we delineated 6,636 distinct retail boundaries that were validated against existing reference retail boundaries. These newly delineated retail boundaries provide valuable insights for researchers, governments, and other relevant stakeholders by enhancing their understanding of retail geography. This dataset can be used as a foundational resource for analyses on topics such as pandemic recovery, retail gentrification, and the resilience of retail spaces in response to e-commerce growth, ultimately contributing to more robust retail sector research in South Korea.

97 MATHEMATICS AND COMPUTING↗

Aircraft Engine Run-To-Failure Data Set Under Real Flight Conditions

The generation of data-driven prognostics models requires the availability of datasets with run-to-failure trajectories. In order to contribute to the development of these methods, the dataset provides a new realistic dataset of run-to-failure trajectories for a small fleet of aircraft engines under realistic flight conditions. The damage propagation modelling used for the generation of this synthetic dataset builds on the modelling strategy from previous work [1] and incorporates two new levels of fidelity. First, it considers real flight conditions as recorded on board of a commercial jet [2]. Secondly, it extends the degradation modelling by relating the degradation process to the operation history. The dataset was generated with the Commercial Modular Aero-Propulsion System Simulation (C-MAPSS) dynamical model [3]. More details about the generation process can be found in [4].

CMAPSS↗

Second-generation downscaled earth system model data using generative machine learning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. As with the first Sup3rCC version, the data still include temperature, wind speed and direction at multiple heights, pressure, three components of downwelling solar radiation, and relative humidity—all at 4-kilometer (km) hourly resolution over the contiguous United States. This is a 25x spatial enhancement and 24x temporal enhancement of the source 100-km daily-average ESM data. This extension of the Sup3rCC dataset includes data from six ESMs from two shared socioeconomic pathways (SSPs) totaling 400 years of data with multiple future projections of changing meteorological conditions. The scenario selection was based on a structured evaluation of historical ESM skill and comprehensive representation of possible trajectories of future climate change in temperature, humidity, precipitation, solar irradiance, and near-surface wind speeds. The inclusion of multiple future projections is intended to enable users to assess key drivers of un 36 certainty and variability. All data are double-bias corrected, resulting in a product that can be used out-of-the-box for energy system analysis with minimal historical bias. The potential applications of Sup3rCC data extend to various topics in renewable energy resource assessment, energy systems modeling, and grid resilience studies. High-resolution future meteorological projections are critical for evaluating the effects of changing meteorological conditions on renewable energy generation, energy demand, and for optimizing energy storage and grid infrastructure. The 4-km hourly resolution of the downscaled data enables understanding of spatial and temporal variability at the scales necessary for energy system operational planning. In addition, the dataset can support risk assessments by providing detailed information on possible future extreme weather events and long-term meteorological variability at scales relevant to energy infrastructure. By offering an enhanced representation of possible future meteorological conditions, the second-generation Sup3rCC dataset enables more precise modeling of energy resilience and adaptation strategies in response to changing meteorological conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

LandScan mosaic enables high-resolution gridded population estimates with explicit uncertainty

Gridded population datasets represent high-resolution distributions of human occupancy, enabling informed decision-making across a broad range of fields. These data products are valuable for assessing environmental risk, urban development, disaster preparedness and resource allocation—areas where accurate population estimates directly enhance policy effectiveness and optimize resource distribution. Despite the importance of gridded population datasets, traditional population modeling approaches often overlook inherent uncertainties in the estimation process. This limitation can create a false sense of certainty in population estimates, potentially leading to flawed decisions by those who rely on the data. To address this methodological gap, we introduce a probabilistic machine learning modeling framework, LandScan Mosaic, that explicitly incorporates uncertainty into the population modeling process. Our approach systematically quantifies uncertainty in three key modeling parameters of the LandScan HD gridded population dataset: building use types, floor counts, and occupancy rates. By employing Monte Carlo simulations, we propagate these uncertainties through the modeling process, yielding probability distributions of population counts in place of deterministic point estimates. We demonstrate the practical application of this framework in Iloilo City, Philippines, using structured decision-making techniques and our probabilistic estimates to identify and prioritize areas most affected by projected flooding, supporting targeted interventions that address both economic and social risks. In doing so, we propose a population-specific approach for incorporating confidence into structured decision making processes. Through a comparative analysis with conventional deterministic approaches and point estimate approaches, including LandScan HD and WorldPop, we evaluate how the incorporation of machine learning and uncertainty influences decision rankings. This research advances population distribution modeling by offering a robust, quantitative approach that explicitly accounts for uncertainty in the underlying data, along with guidance for how users can apply uncertainty in their decision-making.

Environmental sciences↗

Dynamic Boundary Microgrids Under Privatization Considerations

Microgrids have physical, electrical, and logical (data, network, and ownership) boundaries. To power unserved customer loads during an outage, microgrids can extend the traditional operational boundaries. This can become complex when considering microgrid-to-microgrid (M2M) interactions where sensitive information such as competitive microgrid operational data is not shared. This work proposes an optimization method coordinated between microgrid controllers and distribution management systems that limits data sharing. The method involves a competitive bidding strategy that maximizes unserved load coverage while minimizing resource utilization and sensitive operational data sharing among entities. The work is validated on a two-microgrid system with photovoltaic and energy storage systems and curves of load derived from real world residential buildings datasets. Results show that the proposed method, when applied for three distinct use cases of energy storage sufficiency to cover the predefined boundary and/or the expanded boundary, can successfully select and bid the available load coverage.

Starke, Michael [ORNL] (ORCID:0000000221211195)↗

NOODLES Grid [SWR-25-93]

NOODLES Grid is a real-time visualization server for power system simulation data. It loads precomputed datasets, builds optimized instance renderings, and streams live interactive scenes to connected clients. It is built for high scalability, flexible visualization, and fast interaction. NOODLES is a cross-platform/device/tool protocol for collaborative visualization. NOODLES was Developed at the National Renewable Energy Laboratory (NREL) as a capability of the Insight Center https://www.nrel.gov/computational-science/insight-center.html

Brunhart-Lupo, Nicholas [National Renewable Energy↗

The Importance of Being Adaptable: An Exploration of the Power and Limitations of Domain Adaptation for Simulation-Based Inference with Galaxy Clusters

The application of deep machine learning methods in astronomy has exploded in the last decade, with new models showing remarkably improved performance on benchmark tasks. Not nearly enough attention is given to understanding the models' robustness, especially when the test data are systematically different from the training data, or "out of domain." Domain shift poses a significant challenge for simulation-based inference, where models are trained on simulated data but applied to real observational data. In this paper, we explore domain shift and test domain adaptation methods for a specific scientific case: simulation-based inference for estimating galaxy cluster masses from X-ray profiles. We build datasets to mimic simulation-based inference: a training set from the Magneticum simulation, a scatter-augmented training set to capture uncertainties in scaling relations, and a test set derived from the IllustrisTNG simulation. We demonstrate that the Test Set is out of domain in subtle ways that would be difficult to detect without careful analysis. We apply three deep learning methods: a standard neural network (NN), a neural network trained on the scatter-augmented input catalogs, and a Deep Reconstruction-Regression Network (DRRN), a semi-supervised deep model engineered to address domain shift. Although the NN improves results by 17% in the Training Data, it performs 40% worse on the out-of-domain Test Set. Surprisingly, the Scatter-Augmented Neural Network (SANN) performs similarly. While the DRRN is successful in mapping the training and Test Data onto the same latent space, it consistently underperforms compared to a straightforward Yx scaling relation. These results serve as a warning that simulation-based inference must be handled with extreme care, as subtle differences between training simulations and observational data can lead to unforeseen biases creeping into the results.

Ntampaka, Michelle [Baltimore, Space Telescope Sci↗