Search NASA⌕ Search

SEARCH · Search NASA

Results for “dataset use case study”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Case Study: NREL Campus Chilled Water Storage Potential: Benchmark Datasets Development and Applications, Task 4 - Use Case Demonstration

The Benchmark Datasets Development and Applications project is a three-year collaboration between the National Renewable Energy Laboratory (NREL), Oak Ridge National Laboratory, Pacific Northwest National Laboratory, and Lawrence Berkeley National Laboratory. The project seeks to collect and curate high-resolution, well-calibrated time series of building operational and indoor/outdoor environmental data, which are crucial to understanding and optimizing building energy efficiency performance and demand flexibility capabilities as well as benchmarking energy algorithms. Project outcomes include approximately twelve high-fidelity building datasets, enhanced data representation tools, and four case studies to illustrate example applications. The goal of these case studies is to define and execute analyses that demonstrate how one or more datasets collected through this project can address a data gap or challenge historically faced by building stakeholders. This technical paper summarizes the findings of one of these case studies, in which we studied the operational efficiencies of the central cooling system at NREL. We looked at three years of data from the three chillers in the Field Test Laboratory Building (FTLB), from 2019 to 2021, to compare equipment operation and demand throughout the time period. Our analysis indicates that all three chillers are operating at or below the optimal loading conditions for most of the operation time, and thus there was no efficiency drop due to loading of the chillers at full capacity. Our recommendation is that no chiller capacity increase is needed; instead, the central plant could benefit from adopting advanced control logics for optimal sequencing of chillers during part load operations. Analysis of adding chilled water thermal storage to the central plant indicated 34% savings in demand cost and 24.5% savings in total cost (energy consumption and demand charge cost). The payback period is estimated to be 11-22 years with an assumed TES cost of $\$$100-$200 per ton. This case study shows how a selected dataset is used to solve a practical building problem - learning the operational status of its components, analyzing the effectiveness of a proposed new technique, and aiding decision-making for the building operations and maintenance team.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Extract useful information from building permits data to profile a city’s building retrofit history

Building retrofit is one of the key strategies for cities to reduce energy use and GHG emissions. The historical information about changes to buildings is crucial to infer the buildings’ current energy system efficiency levels and to identify candidate buildings for retrofit. In general, a building permit is required before the start of any construction activity of a building, such as changing building structure, remodeling, or installing new equipment. Moreover, many large cities provide public datasets of building permits in history. Therefore, the permits are a potentially good resource for mining information on the city’s retrofit history. In this study, we use the permit dataset from the city of San Francisco as a case study. Location and time information from the dataset is also used to depict the retrofit timeline of each building and the whole building stock. The type of work of the permit is inferred from the descriptive text by a machine learning model. At last, the limitations of the current permit dataset and potential improvements on the permit data management are discussed to better utilize the information in the future.

Zhang, Wanni↗

A Centralized AI Lakehouse Framework for Brain Tumor MRI Classification and Segmentation, University KPI Forecasting, and Water Potability Prediction

In many university and healthcare projects, models are built for very different data types such as tables, institutional time series, and medical images, but they are deployed as separate applications. In this work, that separation made testing and maintenance difficult because each module had its own pipeline and runtime requirements. This paper presents an integrated AI lakehouse-style implementation that runs three model pipelines inside one containerized backend. For medical imaging, we used MRI datasets from IEEE DataPort: a four-class classification set with 7012 images (5708 train/1304 test) and a segmentation set with 3063 image–mask pairs. The classification model (ResNet50 transfer learning) is evaluated using a proper train–validation–test protocol across multiple splits (80/10/10, 70/10/20, 60/10/30, and 10/30/60), achieving a test accuracy of 99.00% under the standard 80/10/10 split. Additionally, a patient-level evaluation is conducted using an external glioma dataset to provide a more realistic assessment without data leakage. The segmentation model (DeepLabV3-ResNet50) achieved 83.09% validation mIoU and 88.79% Dice score. For university KPI forecasting, we used annual IPEDS and NSF HERD data from 2010 to 2023 for three universities (BSU, EOU, and UAB). To examine the effect of preprocessing on forecasting performance, two case studies are conducted. In the first case, linear interpolation is applied to generate semester-level data. In the second case, the original annual data is used directly without interpolation. Random Forest regression and ARIMA models are evaluated using MAE, RMSE, MAPE, and R 2 . The results showed that interpolation improved apparent forecasting performance due to smoothing, while evaluation on the original annual data provided a more realistic assessment of model behavior. To further validate the framework on a larger dataset, an additional case study is conducted using a student dropout dataset. For water potability, we trained and compared multiple tabular classifiers on a large dataset (1,048,575 samples). A Random Forest model (100 trees, max depth 10) achieved 85.86% test accuracy and high recall for unsafe samples (0.8447). All modules are served via FastAPI and deployed together using Docker, with workflow automation routing requests to the correct endpoint. System-level benchmarking indicates that the backend maintains stable throughput and latency under concurrent requests.

97 MATHEMATICS AND COMPUTING↗

Discriminative versus generative approaches to simulation-based inference

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

high energy physics↗

Leveraging Open-Source Satellite-Derived Building Footprints for Height Inference

At a global scale, cities are growing and characterizing the built environment is essential for deeper understanding of human population patterns, urban development, energy usage, climate change impacts, among others. Buildings are a key component of the built environment and significant progress has been made in recent years to scale building footprint extractions from satellite datum and other remotely sensed products. Billions of building footprints have recently been released by companies such as Microsoft and Google at a global scale. However, research has shown that depending on the methods leveraged to produce a footprint dataset, discrepancies can arise in both the number and shape of footprints produced. Therefore, each footprint dataset should be examined and used on a case-by-case study. In this work, we find through two experiments on Oak Ridge National Laboratory and Microsoft footprints within the same geographic extent that our approach of inferring height from footprint morphology features is source agnostic. Regardless of the differences associated with the methods used to produce a building footprint dataset, our approach of inferring height was able to overcome these discrepancies between the products and generalize, as evidenced by 98% of our results being within 3m of the ground-truthed height. This signifies that our approach can be applied to the billions of open-source footprints which are freely available to infer height, a key building metric. This work impacts the broader domain of urban science in which building height is a key, and limiting factor.

Stipek, Clinton [ORNL] (ORCID:0000000280501096)↗

Mapping use cases and dataset needs for benchmarking buildings data

A perennial challenge in buildings research is the lack of high-quality datasets that can be relied upon for a wide array of tasks, including model calibration and improving energy efficiency and load flexibility. Instrumenting a building for data collection is resource intensive, so it is important to be methodical in the approach and ensure that resulting data are flexible and useful for a broad range of analyses. This study aims to fill the gaps in characterizing potential use cases for buildings datasets and mapping them to dataset needs using a well-defined data infrastructure. Here, we have developed a systematic mapping strategy between buildings dataset needs and use cases to help streamline the processes of efficiently targeting datasets, designing building sensing systems, and determining buildings research use cases. We selected 14 prospective use cases and 11 refined buildings data categories for developing the preliminary dataset-needs-to-use-cases mapping matrix (‘DN-UC mapping matrix’) with generic ‘Tags’—a detailed sub-level of data categories extracted by justifying the needs of an aspect of the datasets to use cases. We present two example applications of the developed mapping matrix to demonstrate use of the mapping matrix and its effectiveness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hurricane wind field representation shapes storm surge and building-scale flood hazard estimates

Coastal flood hazard estimates rely on precise hurricane wind forecasts to assess damage and risk. Here, we demonstrate that errors in hurricane wind field representation can lead to significant biases in storm surge and property-level damage estimates. Using Hurricane Ian (2022) as a case study, we compare widely used parametric, reanalysis, and hybrid wind datasets. Improved wind field accuracy reduces storm surge and damage estimate bias by up to 70\%. Our results underscore the importance of accurately predicting hurricane wind structure in hazard assessments.

Coastal Flooding↗

A machine learning method of modern urban building energy modeling: A case study of Chicago

Urban-scale building energy modeling is vital for urban planning. However, it can be challenging to assimilate reliable non-geometry building data for urban-scale modeling without extensive investment. Here, this study introduces a novel approach to developing modern urban-scale building energy stock data using geographic information systems and machine learning algorithms without necessarily requiring pre-supplied non-geometric metadata. The proposed framework integrates building footprint and height data to estimate gross floor areas, and matches each building to a pool of candidate records from ComStock or ResStock—filtered to the same county and ranked by geometric similarity—demonstrate a proof-of-concept case study in Chicago for predicting energy use intensity (EUI) using scalable datasets. The model achieved a mean bias error (MBE) of 0.08 kWh/m² and root mean square error (RMSE) of 14.84 kWh/m² under full metadata input for EUI prediction. With only location inputs, the model captured 69.2 % of EUI within predicted ranges. These results demonstrate the model’s potential to support early-stage urban planning, identify candidates for energy-efficient retrofits. By removing the dependency on detailed pre-surveys or extensive building metadata, the approach overcomes a key barrier in traditional urban-scale building energy modeling, illustrating a pathway toward broader and more cost-effective application, though further multi-city validation and improved treatment of pre-1925 buildings are needed.

Energy Use Intensity↗

PowerModel-AI: A First On-the-Fly Machine-Learning Predictor for AC Power Flow Solutions

The real-time creation of machine-learning models via active or on-the-fly learning has attracted considerable interest across various scientific and engineering disciplines. These algorithms enable machines to build models autonomously while remaining operational. Through a series of query strategies, the machine can evaluate whether newly encountered data fall outside the scope of the existing training set. In this study, we introduce PowerModel-AI, an end-to-end machine learning software designed to accurately predict AC power flow solutions. We present detailed justifications for our model design choices and demonstrate that selecting the right input features effectively captures load flow decoupling inherent in power flow equations. Our approach incorporates on-the-fly learning, where power flow calculations are initiated only when the machine detects a need to improve the dataset in regions where the model’s suboptimal performance is based on specific criteria. Otherwise, the existing model is used for power flow predictions. This study includes analyses of five Texas A&M synthetic power grid cases, encompassing the 14-, 30-, 37-, 200-, and 500-bus systems. The training and test datasets were generated using PowerModels.jl, an open-source power flow solver/optimizer developed at Los Alamos National Laboratory, NM, USA.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Framework to Analyze the Requirements of a Multiport Megawatt-Level Charging Station for Heavy-Duty Electric Vehicles

Widespread adoption of heavy-duty (HD) electric vehicles (EVs) will soon necessitate the use of megawatt (MW)-scale charging stations to charge high-capacity HD EV battery packs. Such a station design needs to anticipate possible station traffic, average and peak power demand, and charging/wait time targets to improve throughput and maximize revenue-generating operations. High-power direct current charging is an attractive candidate for MW-scale charging stations at the time of this study, but there are no precedents for such a station design for HD vehicles. We present a modeling and data analysis framework to elucidate the dependencies of a MW-scale station operation on vehicle traffic data and station design parameters and how that impacts vehicle electrification. This framework integrates an agent-based charging station model with vehicle schedules obtained through real-world vehicle telemetry data analysis to explore the station design and operation space. A case study applies this framework to a Class 8 vehicle telemetry dataset and uses Monte Carlo simulations to explore various design considerations for MW-scale charging stations and EV battery technologies. The results show a direct correlation between optimal charging station placement and major traffic corridors such as cities with ports, e.g., Los Angeles and Oakland. Corresponding parametric sweeps reveal that while good quality of service can be achieved with a mix of 1.2-megawatt and 100-kilowatt chargers, the resultant fast charging time of 35–40 min will need higher charging power to reach parity with refueling times.

33 ADVANCED PROPULSION SYSTEMS↗

Design and Development of a High Fidelity Cyber-Physical Testbed

In order to ensure that future critical infrastructure systems are resilient to various types of such advanced and persistent threats, it is important to develop and integrate tailored solutions that holistically address cyber-attack detection and mitigation in a timely manner such that adverse system impacts that impact a large population are avoided. Further, it is essential to create environments that allow control, protection and communication to exist within a realistic environment to analyze the effects of adverse conditions and system operating modes. This project aims to establish a high-fidelity testbed environment for modeling and simulating a single microgrid all the way up to a network of microgrids along with baseline controls, protection, and associated cyber communication. This is an important activity because accurately modeling and simulating the various power-electronics-based DERs and loads in a microgrid is critical to adequately capturing their behaviors over a wide range of off-normal conditions, as well as to evaluate the resilience of the system using the developed controls. The work presented in this report focuses on the process of building this high-fidelity testbed and the associated experimentation it enables. The model enables the creation of high-fidelity use cases and associated datasets that have been used extensively within the initiative to study resilience and support novel control development and prototyping. The work heavily leverages existing capability that is part of the high-fidelity cyber-physical system experimentation lab to create a power hardware-in-the-loop setup. The report also details the creation of an automated model building platform that can enable high-fidelity real-time models to be built without much effort allowing existing low-fidelity models to be analyzed in higher fidelity. Lastly, the report also discusses efforts center around scaling to large complex power system models to make the experimentation more effective.

97 MATHEMATICS AND COMPUTING↗

Large-scale analysis of structural brain asymmetries in schizophrenia via the ENIGMA consortium

Left–right asymmetry is an important organizing feature of the healthy brain that may be altered in schizophrenia, but most studies have used relatively small samples and heterogeneous approaches, resulting in equivocal findings. We carried out the largest case–control study of structural brain asymmetries in schizophrenia, with MRI data from 5,080 affected individuals and 6,015 controls across 46 datasets, using a single image analysis protocol. Asymmetry indexes were calculated for global and regional cortical thickness, surface area, and subcortical volume measures. Differences of asymmetry were calculated between affected individuals and controls per dataset, and effect sizes were meta-analyzed across datasets. Small average case–control differences were observed for thickness asymmetries of the rostral anterior cingulate and the middle temporal gyrus, both driven by thinner left-hemispheric cortices in schizophrenia. Analyses of these asymmetries with respect to the use of antipsychotic medication and other clinical variables did not show any significant associations. Assessment of age- and sex-specific effects revealed a stronger average leftward asymmetry of pallidum volume between older cases and controls. Case–control differences in a multivariate context were assessed in a subset of the data (N = 2,029), which revealed that 7% of the variance across all structural asymmetries was explained by case–control status. Subtle case–control differences of brain macrostructural asymmetry may reflect differences at the molecular, cytoarchitectonic, or circuit levels that have functional relevance for the disorder. Reduced left middle temporal cortical thickness is consistent with altered left-hemisphere language network organization in schizophrenia.

60 APPLIED LIFE SCIENCES↗

Enhanced descriptor identification and mechanism understanding for catalytic activity using a data-driven framework: revealing the importance of interactions between elementary steps

We report accurate identification of descriptors for catalytic activities has long been essential to the in-depth understanding of catalysis and recently to set the basis for catalyst screening. However, commonly used methods suffer from low accuracy in predictability. This study reports an enhanced approach to accurately identify the descriptors from a kinetic dataset using a machine learning (ML) surrogate model. CO hydrogenation to methanol over Cu-based catalysts was taken as a case study. Our model captures not only the contribution from individual elementary steps but also the interaction between relevant steps within a reaction network, which was found to be essential for high accuracy. As a result, six effective descriptors are identified, which are accurate enough to ensure the trained gradient boosted regression (GBR) model for good prediction of the methanol turnover frequency (TOF) over metal (M)-doped Cu(111) model surfaces (M = Au, Cu, Pd, Pt, Ni). More importantly, going beyond the purely mathematical ML model, the catalytic role of each identified descriptor can be revealed by using model-agnostic interpretation tools, which enhances the insight into the promoting effect of alloying. The trained GBR model outperforms the conventional derivative-based methods in terms of both the predictability and the mechanism understanding. It opens alternative possibilities toward accurate descriptor-based rational catalyst optimization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comprehensive analysis of common polymers using hyphenated TGA-FTIR-GC/MS and Raman spectroscopy towards a database for micro- and nanoplastics identification, characterization, and quantitation

Environmental contamination by micro- and nanoplastics (MNPs) is well documented with potential for their increased accumulation globally. Growing public concern over environmental, ecological, and human exposure to MNPs has led to exponential increase in publications, news articles, and reports. Significant knowledge gap exists in standardized analytical methods for the identification and quantification of MNPs from real world environmental samples. Here, in this study, we report comprehensive datasets utilizing thermogravimetric analyzer (TGA) coupled to a Fourier transformed infrared spectrometer (FTIR) and a gas chromatography/mass spectrometer (GC/MS) with corresponding Raman spectral data for the most common polymers documented to be present in the environment (35 plastics of 12 polymer types), to serve as a base line reference for the identification and quantitation of MNPs. Various parameters for TGA-FTIR-GC/MS data acquisition were optimized. Commercial consumer plastic product compositions were identified using this analytical database. Case studies to showcase the utility of the method for polymer mixtures analysis is included. This dataset would serve towards the development of a collaborative, global, comprehensive, and curated public database for the identification of various MNPs and mixtures.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

UNNT: A novel Utility for comparing Neural Net and Tree-based models

The use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning-enabled computer vision for plant phenotyping: a primer on AI/ML and a case study on stomatal patterning

Abstract Artificial intelligence and machine learning (AI/ML) can be used to automatically analyze large image datasets. One valuable application of this approach is estimation of plant trait data contained within images. Here we review 39 papers that describe the development and/or application of such models for estimation of stomatal traits from epidermal micrographs. In doing so, we hope to provide plant biologists with a foundational understanding of AI/ML and summarize the current capabilities and limitations of published tools. While most models show human-level performance for stomatal density (SD) quantification at superhuman speed, they are often likely to be limited in how broadly they can be applied across phenotypic diversity associated with genetic, environmental, or developmental variation. Other models can make predictions across greater phenotypic diversity and/or additional stomatal/epidermal traits, but require significantly greater time investment to generate ground-truth data. We discuss the challenges and opportunities presented by AI/ML-enabled computer vision analysis, and make recommendations for future work to advance accelerated stomatal phenotyping.

Plant Sciences↗

Analysis of Energy Justice and Equity Impacts from Replacing Peaker Plants with Energy Storage

Transitions to low-carbon energy systems are essential to mitigating and adapting to climate change. Energy storage systems are a key component in achieving a viable decarbonized electric grid. However, decarbonization alone does not guarantee a fairer, more inclusive, or socially just energy system. Energy equity and justice should be integrated in energy system transitions to ensure benefits and burdens are shared equitably. In this paper, we discuss the relationship between energy storage and social equity by assessing the use of energy storage to replace natural gas-fired (NG) peaker plants. Peaker plants are disproportionately located near disadvantaged communities and tend to be older and high emitters of health-affecting fine particulate matter and other pollutants. This paper investigates the equity implications of NG peaker plant replacements with battery energy storage in the context of Washington State’s peaker plants to highlight the human-centered values of retiring the plants. The study performed production cost simulations using the latest Western Electric Coordinating Council Anchor Dataset 2030 case and found that total generation cost, locational marginal price, and total annual emissions were reduced with the replacements. These reductions will have equity benefits on local communities including access to clean air, enhanced health outcomes, and energy burden reductions.

Tarekegne, Bethel W.↗