Search NASA⌕ Search

SEARCH · Search NASA

Results for “Grid Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

NARUC grid data sharing playbook

In 2022, the National Association of Regulatory Utility Commissioners (NARUC) launched an initiative to support its members in addressing issues related to grid data sharing. The Grid Data Sharing Collaborative was funded by the DOE’s Office of Electricity and Office of Cybersecurity, Energy Security, and Emergency Response (CESER). NARUC invited programmatic, policy, technical, and cybersecurity subject matter experts from public utility commissions, utilities, non-governmental organizations, energy service companies, and DOE to join the two-year Grid Data Sharing Collaborative to help develop a flexible framework for states to use as a starting point when navigating complex decision-making inherent in grid data sharing. The framework took shape through a series of intensive workshops during which Collaborative participants explored illustrative use cases to identify data needs, articulate the benefits and risks of sharing such data, and assess the trade-offs. Along the way, participants offered suggestions for how the framework could be used in practice. The purpose of this playbook is to describe the elements of the Grid Data Sharing Framework and to begin supporting its implementation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Grid Data Sharing: Brief Summary of Current State Practices

This brief summary discusses general trends in regulatory approaches to grid data sharing and outlines key areas of alignment and divergence; a summary table at the end of the document outlines the approaches of different states and utilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distribution Grid Model Publication Investigation

Interest in the external exchange of distribution grid model data is growing around the world, driven largely by the challenges and opportunities presented by the increasing amount of generation, storage, and flexible load being embedded within the distribution grid. This report provides an overview of the current state of distribution grid model data sharing, with a focus on the industry-leading activities currently underway in Great Britain (GB). A second report will explore opportunities for external distribution grid model sharing in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Differential Privacy in Grid Kitchen: Implementation & Software Documentation

Sharing of power grid feeder models faces significant challenges due to the potential risk of exposing sensitive operational information. Traditional anonymization techniques have shown notable limitations in other sensitive domains, as evidenced by documented re-identification attacks that combine supposedly anonymized datasets with auxiliary information, raising concerns that similar vulnerabilities could affect power grid data. Consequently, there is a pressing need for a more rigorous privacy protection strategy that not only delivers formal mathematical guarantees but also preserves the analytical value of the shared models. To address this challenge, we have enhanced the Grid Kitchen framework by implementing differential privacy mechanisms within the distribution model dehydration pipeline. This implementation carefully calibrates and applies noise to sensitive attributes in feeder models according to configurable privacy levels—low, moderate, and high—each offering different balances between data utility and privacy protection. Our approach uses established noise functions (Gaussian for continuous data and Discrete Laplace for integer values) with parameters carefully calibrated so that the impact of individual data points is effectively masked in the final output. The integration leverages our Noise Catalog, which we developed to categorize feeder model properties by component type, data type, and sensitivity. This catalog guides the application of appropriate noise functions and privacy parameters ($\varepsilon$ and $\delta$) to each attribute, ensuring consistent privacy protection across the model while maintaining its structural integrity and analytical usefulness. This implementation also includes evaluation tools that allow model owners to assess the impact of privacy-preserving transformations before sharing data with external parties. This report provides documentation for the differential privacy capabilities added to the Grid Kitchen project. It includes a primer on differential privacy concepts and their importance in modern data sharing, details the architecture of our implementation, explains the privacy modes and parameter configurations, and offers practical guidance on using the code for applying differential privacy to grid feeder models. Through examples and code snippets, we demonstrate the effective application of these privacy-enhancing technologies, enabling utility operators and researchers to confidently share grid data while protecting sensitive information.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data Request for the Distribution Grid Atlas

Model-based, distribution powerflow analysis is a foundational component of system planning and grid modernization efforts, but data security is an impediment to collaboration among utility engineers, researchers, developers, community members, and other stakeholders. Pacific Northwest National Laboratory (PNNL) and the National Renewable Energy Laboratory (NREL) are partnering to develop the new Distribution Grid Atlas - a publicly available catalog of realistic, geographically relevant, representative distribution feeder models without sensitive geographic information, customer data, or disclosure of utility models. We are looking for utilities to share data for the Distribution Grid Atlas.

distribution↗

Mitigating Data Center Impact on Grid Stability: A Coordinated Control Strategy Using Verrus StabiliGrid Architecture

Large data centers, which now represent a significant and growing share of the total U.S. grid load, can inadvertently destabilize the electrical grid when they disconnect simultaneously during brief voltage disturbances. The July 10, 2024, Eastern Interconnection incident, in which a sub-100-millisecond transmission fault triggered the cascading loss of approximately 1,500 MW of data center load, illustrates this vulnerability. While commercial battery energy storage systems (BESS) deployed in data centers provide device-level fault ride-through per IEEE 1547, they lack coordination with facility protection logic and uninterruptible power supplies (UPS), limiting their effectiveness as grid-stabilizing assets. This report presents the Verrus StabiliGrid architecture, a coordinated control framework that integrates BESS, UPS, and point-of-interconnection (POI) protection settings to enable data centers to ride through both undervoltage and overvoltage grid contingencies without disconnecting. The four-step strategy encompasses: (1) high-resolution power quality monitoring to detect the grid state during events such as undervoltage, overvoltage, underfrequency, and overfrequency; (2) POI protection settings that allow for extended ride-through and grid-connected operation during grid contingencies; (3) grid state-driven autonomous dispatch of assets to improve grid resilience by reducing power draw during undervoltage or absorbing more power during overvoltage events; and (4) coordinated post-recovery dispatch of data center assets to restore firm load to pre-contingency levels. Validated through controller-hardware-in-the-loop (C-HIL) simulations at the National Laboratory of the Rockies, results show grid import restoration to pre-fault levels within 100 milliseconds of voltage recovery. This work advances the ability of data centers to transition from passive, disturbance-sensitive loads to active participants in grid stability, a capability increasingly required by emerging NERC and ERCOT regulatory frameworks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Super-Resolution for Renewable Energy Resource Data with Wind from Reanalysis Data and Application to Ukraine

With a potentially increasing share of the electricity grid relying on wind to provide generating capacity and energy, there is an expanding global need for historically accurate, spatiotemporally continuous, high-resolution wind data. Conventional downscaling methods for generating these data based on numerical weather prediction have a high computational burden and require extensive tuning for historical accuracy. In this work, we present a novel deep learning-based spatiotemporal downscaling method using generative adversarial networks (GANs) for generating historically accurate high-resolution wind resource data from the European Centre for Medium-Range Weather Forecasting Reanalysis version 5 data (ERA5). In contrast to previous approaches, which used coarsened high-resolution data as low-resolution training data, we use true low-resolution simulation outputs. We show that by training a GAN model with ERA5 as the low-resolution input and Wind Integration National Dataset Toolkit (WTK) data as the high-resolution target, we achieved results comparable in historical accuracy and spatiotemporal variability to conventional dynamical downscaling. This GAN-based downscaling method additionally reduces computational costs over dynamical downscaling by two orders of magnitude. We applied this approach to downscale 30 km, hourly ERA5 data to 2 km, 5 min wind data for January 2000 through December 2023 at multiple hub heights over Ukraine, Moldova, and part of Romania. With WTK coverage limited to North America from 2007–2013, this is a significant spatiotemporal generalization. The geographic extent centered on Ukraine was motivated by stakeholders and energy-planning needs to rebuild the Ukrainian power grid in a decentralized manner. This 24-year data record is the first member of the super-resolution for renewable energy resource data with wind from the reanalysis data dataset (Sup3rWind).

17 WIND ENERGY↗

TrustDER: Trusted, Private and Scalable Coordination of Distributed Energy Resources

In this project, the Stanford and SLAC Teams have developed a Trusted, Private and Scalable platform for coordinating Coordination of Distributed Energy Resources (TrustDER). This is a layered system that ensures private, trusted and scalable coordination and monitoring of DERs. It accommodates a variety of resources, such as solar generation, gensets and loads, with a particular focus on battery systems-based resources, as they are a transformational technology experiencing fast growth in adoption by large critical facilities. The platform can be used as standalone or added to existing aggregation systems to enable trust, privacy and resilience. TrustDER consists of layers that address each of the shortcomings of the existing state of the art. Each layer in the platform can operate independently but provides information to the layers above it to enable a novel form of overall coordination architecture. The project consists of several tasks, with each task dedicated to the design of each layer. Task 2 Resource Virtualization defined a software abstraction layer for distributed energy resources (DERs). The goal of this abstraction was to simplify the implementation of algorithms utilizing cooperation of DERs resources in a variety of use cases. Task 3 is on Secure ID for Asset Authentication. Identity Management Systems (IDMS) are a foundational infrastructure for interactions between entities (organizations, users, devices, and services). Secure ID is blockchain-based a distributed identity management system allowing (1) identity provisioning, (2) authentication, (3) authorization, and (4) identity data sharing for IoT-enabled assets on the electricity grid. In this project, the SLAC team focused on designing and testing Keymaker, a protocol for authenticating device identity managed by Secure ID. Task 5 Private and Safe Integration is focused on the design and evaluation of a DER cooperation scheme which allows for the aggregation of DERs without impacting network reliability. The approach is designed based on realistic assumptions regarding data availability, communication infrastructure limitations, and privacy. Task 6 Scalable Distributed Privacy for Information explored how virtualized batteries could be managed privately. Specifically, it examined the case in which a principal provides a partitioned battery to multiple clients. Task 7 Use Cases was to ensure that this technology was applied in relevant situations and scenarios. Primarily, this means that virtualization needed to be employed in a manner that either improved flexibility, bolstered security or privacy, or decreased costs.

25 ENERGY STORAGE↗

Workshop Summary: Bridging the Gap Between Atmospheric Science and Grid Integration

The need for dedicated, accurate, expertly curated weather data is increasingly important as the share of variable renewable energy increases on the power system. Projections for futures with very high (50+% annual energy) shares of variable generation require ongoing assessment of data requirements from industry stakeholders in their power system operation and planning contexts. In March 2024, NREL organized a workshop entitled "Bridging the Gap Between Atmospheric Science and Grid Integration Workshop", which brought atmospheric scientists and power system experts together to refine the requirements of atmospheric datasets for grid integration, and to describe a holistic approach to creating new and regularly updated national scale wind datasets for power system planning and operations. The results of this workshop are being used to inform the near-term development and a longer-term strategy for DOE to produce relevant wind resource datasets and inform wider use of wind/solar/load data sets in power system planning. This presentation provides an overview of a preworkshop survey, an assessment of current state of the art of national-scale datasets for wind resource assessment and grid integration, insights on appropriate uses of the WTK-LED, power system perspectives on data needs, as well as recommended next steps as discussed in the workshop and how these steps support longer-term strategies.

17 WIND ENERGY↗

Federated Machine Learning-Based Anomaly Detection System for Synchrophasor Network Using Heterogeneous Data Sets: Preprint

Synchrophasor technology is widely deployed in the energy management system to monitor the grid health at micro level and perform necessary corrective actions in real time; however, integrated phasor devices and data aggregators are exposed to several cybersecurity threats. This paper proposes a federated ML(FML)-based ADS to detect several data integrity attacks in the synchrophasor network. The proposed approach integrates the horizontal FML technique and consists of substation-based local models and a control center-based global model. The proposed methodology includes training local models using heterogeneous data sets that include network and grid information and updating the global model through multiple iterations by sharing model gradients. Finally, the trained global model is applied to identify cyberattacks, normal operation, and physical events. To validate the proof of concept, we used synthetic data sets generated by Mississippi State University and Oak Ridge National Laboratory for training and testing the classification models using the National Renewable Energy Laboratory's high performance computing resources. Our experimental results, computed through several performance measures, reveal that the proposed approach shows consistent performance during the binary, three-class, and multiclass classifications while ensuring privacy of synchrophasor data.

anomaly detection system↗

County Level Annual Population Projections for SSP3 and SSP5

This dataset consist of annual (2020-2100) county-level population projections for the United States (U.S.) for Shared Socioeconomic Pathways (SSPs) 3 and 5. The original decadal state-level data is used as an input to the gridded population data model to downscale the state-level projections for each SSP to a 1 km grid. The 1 km data is then aggregated to the county-level using the 2020 TIGER/Line county shapefiles from the U.S. Census Bureau. The two data files are shared as .csv files with the following structure: Rows = Data for individual counties which are identified using their Federal Information Processing Standard (FIPS) code. The FIPS code is stored in the first column. Columns = Each year from 2020-2100 is an individual column. For easier reference, the state name associated with each row is stored in the last column. Values in each cell are the downscaled estimated population for that county-year combination. Values are fractional due to the weighting scheme used in the downscaling. The paper and model source code cited in the "Related Works" section describes the initial source of the population projections.

FIPS↗

Modelling and analysis of nuclear reactor system coupled with a liquid metal battery

Traditionally, nuclear power plants in the U.S. provide baseload power to the power grid because they have less flexibility for ramping their output power than natural gas peaking plants. However, achieving climate goals to reduce the consumption of fossil‐based natural gas places pressure on nuclear power plants and other power generators to ramp up their power output to balance grid generation with demand. This paper presents the modelling and performance analysis of a nuclear reactor system (NRS) coupled to a liquid‐metal battery (LMB) to improve its dynamic response and enable its black start capability. The NRS and LMB thermal behaviour are modelled in Dymola, while the electrical dynamics of the LMB and power grid are modelled in RTDS‐RSCAD. Both simulation platforms are coupled and share their thermal and electrical data using a Transmission Control Protocol/Internet Protocol (TCP/IP) communication protocol. The dynamic performance of the NRS‐LMB integration is tested on the IEEE 9 bus, which demonstrates its ability to respond and provide frequency and voltage regulation. The black start capability of the NRS‐LMB is also evaluated by simulating a grid outage and using the LMB to supply the auxiliary loads required to bring the NRS back online as soon as possible. The results show that coupling an NRS to an LMB improves the system dynamic performance and enables it to black start after being disconnected from the grid for several days.

25 ENERGY STORAGE↗

Hourly Electricity Demand Profiles for Each County in the Contiguous United States

This dataset provides estimated hourly electricity demand for each county in the contiguous United States from 2016-2023. The demand profiles represent the sum of two components: (1) Weighted averages of reported hourly demand profiles for North American Electric Reliability Corporation balancing authority (BA) regions and subregions, scaled to match annual estimates of county-level retail sales and direct use of electricity and weighted by the estimated percentage of county load served by each BA region or subregion. (2) Weighted averages of modeled hourly, county- and sector-level distributed photovoltaic (DPV) capacity factor profiles, scaled to match annual estimates of on-site consumption of DPV-generated electricity for each county and weighted by the percentage of consumption attributable to each sector Annual county-level retail sales are estimated by aggregating utility-reported sales to the state level and allocating the results to counties according to each county's share of state population. Annual county-level direct use is calculated by aggregating power plant-reported direct use values. Annual county-level on-site consumption of DPV-generated electricity is estimated by aggregating utility-reported net metering data to determine the amount of DPV-generated electricity sold back to the grid for each state, subtracting those values from modeled state-level DPV generation estimates, and allocating the results to counties according to each county's share of statewide modeled DPV generation. The open-source Python code used to develop this dataset is available at "Historical Load Data Repository" link below.

14 SOLAR ENERGY↗

Impact of Hydrological Data on Power System Operational Studies: Preprint

Hydropower is expected to play an important role in maintaining grid reliability and flexibility as the share of of variable renewable energy increases. While the current hydropower operational models have been studied and used widely, they haven't been updated for decades to meet new performance standards. For example, current steady state and dynamic models often neglect hydrological conditions, which may lead to unrealistic expectations when relying on hydropower for energy and ancillary services. To study this impact, a multi-timescale hydrological model was created by leveraging the National Renewable Energy Laboratory-developed Multi-timescale Integrated Dynamics and Scheduling (MIDAS) tool. Using MIDAS, we compare the impact of considering hydrological conditions in a day-ahead unit commitment (DAUC) schedule on the reduced 240-bus Western Interconnect (WI) test system under winter and summer case studies. We show that neglecting current hydrological conditions of hydropower plants in power system models can lead to an overestimation of hydropower capabilities, which could lead to power balancing issues. For example, power system DAUC simulation results of our reduced test system show that in the case where hydrological conditions are not considered in the model, an approximate 31% overestimation of hydropower capabilities occurs in the summer case and approximately 60% occurs in the winter case compared to what is available. Additionally, results show an underestimation of WI day-ahead power system generation costs by approximately $54M - $80M in the weekly summer scenario and $116M - $126M in the weekly winter scenario. This analysis helps to underscore the importance of considering hydrological data in power system operational studies.

ENERGY PLANNING, POLICY, AND ECONOMY,HYDRO ENERGY↗

Multi-Agent Graph-Attention Deep Reinforcement Learning for Post-Contingency Grid Emergency Voltage Control

Grid emergency voltage control (GEVC) is paramount in electric power systems to improve voltage stability and prevent cascading outages and blackouts in case of contingencies. While most deep reinforcement learning (DRL)-based paradigms perform single agents in a static environment, real-world agents for GEVC are expected to cooperate in a dynamically shifting grid. Moreover, due to high uncertainties from combinatory natures of various contingencies and load consumption, along with the complexity of dynamic grid operation, the data efficiency and control performance of the existing DRL-based methods are challenged. To address these limitations, we propose a multi-agent graph-attention (GATT)-based DRL algorithm for GEVC in multi-area power systems. Here, we develop graph convolutional network (GCN)-based agents for feature representation of the graph-structured voltages to improve the decision accuracy in a data-efficient manner. Furthermore, a cutting-edge attention mechanism concentrates on effective information sharing among multiple agents, synergizing different-sized subnetworks in the grid for cooperative learning. We address several key challenges in the existing DRL-based GEVC approaches, including low scalability and poor stability against high uncertainties. Test results in the IEEE benchmark system verify the advantages of the proposed method over several recent multi-agent DRL-based algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗