Search NASASearch

SEARCH · Search NASA

Results for “Data modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Review of data-driven models for quantifying load shed by non-residential buildings in the United States

Shifting and shedding power demand in buildings can be cost-effective techniques for grids to function reliably and for end users to earn compensation. Grid operators reimburse customers in proportion to the quantity of load shed. Simple data-driven methods are used to quantify this shed, which is the difference between a measured load during the event and modeled "baseline" that would have occurred in absence of the event. These methods have evolved over the years and in many cases have been integrated with building physics, to make them a hybrid between physics based and empirical models. However, there is no comprehensive analysis that provides guidance to building operators, grid operators and researchers in selecting appropriate models based on their specific needs and available data. Here, this work aims to fill this gap by critically assessing the performance of baseline models put forward from the year 2000 through 2023. The literature reviewed includes reports generated by grid operators, reports from national laboratories and academic journal articles. The work outlines modeling features like the inputs, training period, estimation method, adjustments to fine tune the predictions and metrics to evaluate the performance. A comprehensive list of 50 models has been provided. For each model, the study explores the applicability of the model to weather sensitive buildings, variability in the building profile, timing of the event, and whether the building reduces energy consumption before an event. The work identifies the situations in which a particular model works and draws lessons based on evidence of performance. Finally, recommendations to aid in model selection are given.

97 MATHEMATICS AND COMPUTING

A 30-yr high-resolution weather research and forecasting model downscaling data over California and Nevada

This dataset presents a 30-year high resolution meteorological dataset obtained using the WRF model (Advanced version Research WRF version 4.4). We used WRF and European Centre for Medium-Range Weather Forecasts Reanalysis v5 as initial and boundary conditions to generate gridded meteorological variables. A large number of surface weather stations was used for model validation. A multi-physics analysis was first developed to identify a good physics suite extended from 6 November 00 UTC to 10 November 23 UTC, 2018, which included the Camp Fire in northern California. Based on the best physics suite, the downscaling dataset extends from 1 December to 28 February, 1990–2021 and the horizontal domain has 1.5 km grid spacing covering the entire states of California and Nevada in the United States. Comparisons between hourly surface observations and WRF simulations of air temperature, relative humidity and wind speeds show mean absolute errors on the order of (1.6-2.0 C), (10 %) and 1.2–1.5 m s -1 , respectively.

54 ENVIRONMENTAL SCIENCES

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES

Conservative projection-based data-driven model order reduction of a fluid-kinetic spectral solver

Kinetic simulations are computationally intensive due to six-dimensional phase space discretization. Many kinetic spectral solvers use the asymmetrically weighted Hermite expansion due to its conservation and fluid-kinetic coupling properties, i.e., the lower-order Hermite moments capture and describe the macroscopic fluid dynamics, and higher-order Hermite moments describe the microscopic kinetic dynamics. We leverage this structure by developing a parametric data-driven reduced-order model based on the proper orthogonal decomposition, which projects the higher-order kinetic moments while retaining the fluid moments intact. We demonstrate analytically and numerically that the method ensures local and global mass, momentum, and energy conservation. The numerical results show that the proposed method effectively replicates the high-dimensional spectral simulations at a fraction of the computational cost and memory, as validated on the weak Landau damping and two-stream instability benchmark problems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING

Improved Muon Energy Estimation Using a Detailed Model of Multiple Coulomb Scattering in the MicroBooNE LArTPC

We present an improved technique for estimating a muon's energy by measuring the deflections along its path inside the MicroBooNE detector from multiple Coulomb scattering (MCS). This approach implements several innovations that better capture detector non-idealizations compared to previous MCS-based muon energy estimators. As a result, it achieves improved resolution, reduced bias, and better data-model agreement. Using model simulation, for fully contained events the estimated bias is within 1\% and the estimated resolution narrows from 10\% to 4.3\% as muon energy increases from 0.1\,GeV to 2\,GeV. For events with particles exiting the detector volume, at least a meter of reconstructed muon track, and a muon energy below 2\,GeV, the estimated bias is less than 2\% and the estimated resolution varies from 7\% to 17\% over muon energy. These demonstrate significant improvements over the performance of previous work using an MCS-based energy estimator at MicroBooNE~\cite{mcs_2017}, which exhibited approximately twice worse resolution and a bias of 20\% over the same energy region. Data-model goodness-of-fit studies are used to validate the estimator's performance on data, showing good agreement within model uncertainties.

Cooper-Troendle, London [U. Pittsburgh (main); Fer

Forecasting Solar Photovoltaic Power Production: A Comprehensive Review and Innovative Data-Driven Modeling Framework

The intermittent and stochastic nature of Renewable Energy Sources (RESs) necessitates accurate power production prediction for effective scheduling and grid management. This paper presents a comprehensive review conducted with reference to a pioneering, comprehensive, and data-driven framework proposed for solar Photovoltaic (PV) power generation prediction. The systematic and integrating framework comprises three main phases carried out by seven main comprehensive modules for addressing numerous practical difficulties of the prediction task: phase I handles the aspects related to data acquisition (module 1) and manipulation (module 2) in preparation for the development of the prediction scheme; phase II tackles the aspects associated with the development of the prediction model (module 3) and the assessment of its accuracy (module 4), including the quantification of the uncertainty (module 5); and phase III evolves towards enhancing the prediction accuracy by incorporating aspects of context change detection (module 6) and incremental learning when new data become available (module 7). This framework adeptly addresses all facets of solar PV power production prediction, bridging existing gaps and offering a comprehensive solution to inherent challenges. By seamlessly integrating these elements, our approach stands as a robust and versatile tool for enhancing the precision of solar PV power prediction in real-world applications.

14 SOLAR ENERGY

A New Coupled Biogeochemical Modeling Approach Provides Accurate Predictions of Methane and Carbon Dioxide Fluxes Across Diverse Tidal Wetlands

Abstract Tidal wetlands provide valuable ecosystem services, including storing large amounts of carbon. However, the net exchanges of carbon dioxide (CO 2 ) and methane (CH 4 ) in tidal wetlands are highly uncertain. While several biogeochemical models can operate in tidal wetlands, they have yet to be parameterized and validated against high‐frequency, ecosystem‐scale CO 2 and CH 4 flux measurements across diverse sites. We paired the Cohort Marsh Equilibrium Model (CMEM) with a version of the PEPRMT model called PEPRMT‐Tidal, which considers the effects of water table height, sulfate, and nitrate availability on CO 2 and CH 4 emissions. Using a model‐data fusion approach, we parameterized the model with three sites and validated it with two independent sites, with representation from the three marine coasts of North America. Gross primary productivity (GPP) and ecosystem respiration (R eco ) modules explained, on average, 73% of the variation in CO 2 exchange with low model error (normalized root mean square error (nRMSE) <1). The CH 4 module also explained the majority of variance in CH 4 emissions in validation sites ( R 2 = 0.54; nRMSE = 1.15). The PEPRMT‐Tidal‐CMEM model coupling is a key advance toward constraining estimates of greenhouse gas emissions across diverse North American tidal wetlands. Further analyses of model error and case studies during changing salinity conditions guide future modeling efforts regarding four main processes: (a) the influence of salinity and nitrate on GPP, (b) the influence of laterally transported dissolved inorganic C on R eco , (c) heterogeneous sulfate availability and methylotrophic methanogenesis impacts on surface CH 4 emissions, and (d) CH 4 responses to non‐periodic changes in salinity.

54 ENVIRONMENTAL SCIENCES

Information theory optimization of signals from small-angle scattering measurements

Small-angle X-ray scattering (SAXS) of particles in solution informs on the conformational states and assemblies of biological macromolecules (bioSAXS) outside of cryo- and solid-state conditions. In bioSAXS, the SAXS measurement under dilute conditions is resolution limited, and through an inverse Fourier transform, the measured SAXS intensities directly relate to the physical space occupied by the particles via the P (r)-distribution. Yet, this inverse transform of SAXS data has been historically cast as an ill-posed, ill-conditioned problem requiring an indirect approach. Here, we show that through the applications of matrix and information theories, the inverse transform of SAXS intensity data is a well-conditioned problem. The so-called ill-conditioning of the inverse problem is directly related to the Shannon number. By exploiting the oversampling enabled by modern detectors, a direct inverse Fourier transform of the SAXS data is possible, provided the recovered information does not exceed the Shannon number. The Shannon limit corresponds to the maximum number of significant singular values that can be recovered in a SAXS experiment, suggesting this relationship is a fundamental property of band-limited inverse integral transform problems. This correspondence reduces the complexity of the inverse problem to the Shannon limit and maximum dimension. We propose a hybrid scoring function using an information theory framework that assesses both the quality of the model-data fit as well as the quality of the recovered P (r)-distribution. The hybrid score utilizes the Akaike information criteria and Durbin-Watson statistic that considers parameter-model complexity, i.e., degrees of freedom, and the randomness of the model-data residuals. The described tests and findings extend the boundaries for bioSAXS by completing the information theory formalism initiated by Peter B. Moore to enable a quantitative measure of resolution in SAXS, robustly determine maximum dimension, and more precisely define the best parameter model appropriately representing the observed scattering data.

Rambo, Robert P. [Science and Technology Facilitie

Distribution Grid Model Publication Investigation

Interest in the external exchange of distribution grid model data is growing around the world, driven largely by the challenges and opportunities presented by the increasing amount of generation, storage, and flexible load being embedded within the distribution grid. This report provides an overview of the current state of distribution grid model data sharing, with a focus on the industry-leading activities currently underway in Great Britain (GB). A second report will explore opportunities for external distribution grid model sharing in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION

Forecasting for ESCAPE: A Multi-Institution Hybrid Forecasting and Nowcasting Operation for Sea-Breeze Convection Supporting a Ground-Based and Airborne Field Campaign

The Experiment of Sea-Breeze Convection, Aerosols, Precipitation and Environment (ESCAPE) field project deployed two aircraft and ground-based assets in the vicinity of Houston, Texas, between 27 May and 2 July 2022, examining how meteorological conditions, dynamics, and aerosols control the initiation, early growth stage, and evolution of coastal convective clouds. To ensure that airborne- and ground-based assets were deployed appropriately, a forecasting and nowcasting team was formed. Daily forecasts guided real-time decision-making by assessing synoptic weather conditions, environmental aerosol, and a variety of atmospheric modeling data to assign a probability for meeting specific ESCAPE campaign objectives. During the research flights, a small team of forecasters provided “nowcasting” support by analyzing radar, satellite, and new model data in real time. The nowcasting team proved invaluable to the campaign operation, as sometimes changing environmental conditions affected, for example, the timing of convective initiation. In addition to the success of the forecasting and nowcasting teams, the ESCAPE campaign offered a unique “testbed” opportunity where in-person and virtual support both contributed to campaign objectives. The forecasting and nowcasting teams were each composed of new and experienced forecasters alike, where new forecasters were given invaluable experience that would otherwise be difficult to attain. Both teams received training on forecast models, map analysis, Hybrid Single-Particle Lagrangian Integrated Trajectory model (HYSPLIT), and thermodynamic sounding analysis before the beginning of the campaign. In this article, the ESCAPE forecasting and nowcasting teams reflect on these experiences, providing potentially useful advice for future field campaigns requiring forecasting and nowcasting support in a hybrid virtual/in-person framework.

54 ENVIRONMENTAL SCIENCES

Data-driven Modeling for Grid Edge IBRs: A Digital Twin Perspective of User-Defined Models

Recent events in Odessa have brought attention to the challenges associated with the interaction between Inverter- Based Resources (IBRs) and the transmission and distribution system. The NERC event diagnosis report has highlighted sev- eral issues, emphasizing the need for continuous performance monitoring of these IBRs by system operators. Key areas of concern include the mismatch of control and protection perfor- mance of IBRs between the original equipment manufacturer (OEM)-provided models and field measurements. The inability to replicate the realistic response can result in incorrect reliability and resilience studies. In this paper, we developed an approach on how to emulate the behavior of an IBR using measurement data obtained for system operators to utilize in real-time and long- term planning. Two experiments are conducted in the phasor domain and electromagnetic transients (EMT) domain to emulate the behavior for grid forming and grid following inverters under various operating conditions and the effectiveness of the proposed model is demonstrated in terms of accuracy and ease of utilizing user-defined models (UDMs)

Mahapatra, Kaveri [BATTELLE (PACIFIC NW LAB)]

Data-Driven Modeling of High-Resolution Residential Load Profiles Using Low-Resolution Smart Meter Measurements

Accurate and high-resolution residential load profiles are essential for power system modeling, demand response planning, and effective grid operation. As the energy sector moves towards a more actively managed distribution system, the ability to understand residential energy consumption at a minute-by-minute scale becomes increasingly critical. High-resolution load profiles provide key insights into demand patterns and user behavior, enabling grid operators to design more effective energy solutions; however, residential load measurements in the field are typically recorded at low resolutions, such as 15-60 minutes, which makes it hard to study the characteristics of different residential customers. This paper addresses these challenges by introducing a data-driven approach to generate realistic, high-resolution residential load profiles based on lowre-solution measurements and weather information. The proposed method retains the key features of the actual residential load measurements while offering appliance-level energy consumption details for each residential building. The results demonstrate the effectiveness of the proposed load profile generator, proving its capability to support utilities in optimizing residential energy management and ensuring a more reliable and resilient grid.

24 POWER TRANSMISSION AND DISTRIBUTION

A change language for ontologies and knowledge graphs

Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Data-driven modeling of dislocation mobility from atomistics using physics-informed machine learning

Dislocation mobility, which dictates the response of dislocations to an applied stress, is a fundamental property of crystalline materials that governs the evolution of plastic deformation. Traditional approaches for deriving mobility laws rely on phenomenological models of the underlying physics, whose free parameters are in turn fitted to a small number of intuition-driven atomic scale simulations under varying conditions of temperature and stress. This tedious and time-consuming approach becomes particularly cumbersome for materials with complex dependencies on stress, temperature, and local environment, such as body-centered cubic crystals (BCC) metals and alloys. In this paper, we present a novel, uncertainty quantification-driven active learning paradigm for learning dislocation mobility laws from automated high-throughput large-scale molecular dynamics simulations, using Graph Neural Networks (GNN) with a physics-informed architecture. We demonstrate that this Physics-informed Graph Neural Network (PI-GNN) framework captures the underlying physics more accurately compared to existing phenomenological mobility laws in BCC metals.

36 MATERIALS SCIENCE