Search NASA⌕ Search

SEARCH · Search NASA

Results for “Temporal data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA ↗

Temporally-consistent koopman autoencoders for forecasting dynamical systems

Absence of sufficiently high-quality data often poses a key challenge in data-driven modeling of high-dimensional spatio-temporal dynamical systems. Koopman Autoencoders (KAEs) harness the expressivity of deep neural networks (DNNs), the dimension reduction capabilities of autoencoders, and the spectral properties of the Koopman operator to learn a reduced-order feature space with simpler, linear dynamics. However, the effectiveness of KAEs is hindered by limited and noisy training datasets, leading to poor generalizability. To address this, we introduce the Temporally-Consistent Koopman Autoencoder (tcKAE), designed to generate accurate long-term predictions even with limited and noisy training data. This is achieved through a consistency regularization term that enforces prediction coherence across different time steps, thus enhancing the robustness and generalizability of tcKAE over existing models. We provide analytical justification for this approach based on Koopman spectral theory and empirically demonstrate tcKAE’s superior performance over state-of-the-art KAE models across a variety of test cases, including simple pendulum oscillations, kinetic plasma, and fluid flow data.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Experimental data for damage mechanics simulation challenge

While there are many computational approaches for simulating damage in rock and other materials, few have been ground truth tested with either known experimental data or with blind data sets. Here, in this work, we present a bench-mark laboratory data set for a damage mechanics challenge to compare computational approaches on damage evolution in brittle-ductile materials. The samples were fabricated through additive manufacturing to produce repeatable specimens designed to fail in controlled ways. The failure was induced in the samples using a 3-point bending test to produce different Modes such as Mode I and mixed Modes including I-II, I-III and I-II-III Modes to generate a calibration data set and a blind challenge data set. Data collected included spatial and temporal measurements from traditional digital load–displacement sensors, 2D digital image correlation measurement to map surface deformations, 3D X-ray microscopy to ground-truth the crack-failure geometry, and laser profilometry to capture surface roughness. The data sets are available, on a data repository, to the community to advance computational models to improve our ability to predict damage in brittle-ductile materials.

3-point bending↗

SigTime: Learning and Visually Explaining Time Series Signatures

Understanding and distinguishing temporal patterns in time series data is essential for scientific discovery and decision-making. For example, in biomedical research, uncovering meaningful patterns in physiological signals can improve diagnosis, risk assessment, and patient outcomes. However, existing methods for time series pattern discovery face major challenges, including high computational complexity, limited interpretability, and difficulty in capturing meaningful temporal structures. Here, to address these gaps, we introduce a novel learning framework that jointly trains two Transformer models using complementary time series representations: shapelet-based representations to capture localized temporal structures and traditional feature engineering to encode statistical properties. The learned shapelets serve as interpretable signatures that differentiate time series across classification labels. Additionally, we develop a visual analytics system—SigTime—with coordinated views to facilitate exploration of time series signatures from multiple perspectives, aiding in useful insights generation. We quantitatively evaluate our learning framework on eight publicly available datasets and one proprietary clinical dataset. Additionally, we demonstrate the effectiveness of our system through two usage scenarios along with the domain experts: one involving public ECG data and the other focused on preterm labor analysis.

97 MATHEMATICS AND COMPUTING↗

Spatiotemporal Downscaling Model for Solar Irradiance Forecast Using Nearest-Neighbor Random Forest and Gaussian Process

Accurate solar photovoltaic (PV) capacity estimation requires high-resolution, site-specific solar irradiance data to account for localized variability. However, global datasets, such as the National Solar Radiation Database (NSRDB), provide regional averages that fail to capture the fine-scale fluctuations critical for large-scale grid integration. This limitation is particularly relevant in the context of increasing distributed energy resources (DERs) penetration, such as rooftop PV. Additionally, it is critical to the implementation of the U.S. Federal Energy Regulatory Commission (FERC) Order 2222, which facilitates DER participation in U.S. bulk power markets. To address this challenge, this study evaluates Nearest-Neighbor Random Forest (NNRF) and Nearest-Neighbor Gaussian Process (NNGP) models for spatiotemporal downscaling of global solar irradiance data. By leveraging historical irradiance and meteorological data, these models incorporate spatial, temporal, and feature-based correlations to enhance local irradiance predictions. The NNRF model, a machine-learning approach, prioritizes computational efficiency and predictive accuracy, while the NNGP model offers a level of interpretability and prediction uncertainty by numerically quantifying correlations and dependencies in the data. Model validation was conducted using day-ahead predictions. The results showed that the average Goodness of Fit (GoF) of the NNRF model of 90.61% across all eight sites outperformed the GoF of the NNGP of 85.88%. Additionally, the computational speed of NNRF was 2.5 times faster than the NNGP. Finally, the NNGP displayed polynomial scaling while the NNRF scaled linearly with increasing number of nearest neighbors. Additional validation of the model on five sites in Puerto Rico further confirmed the superiority of the NNRF model over the NNGP model. These findings highlight the robustness and computational efficiency of NNRF for large-scale solar irradiance downscaling, making it a strong candidate for improving PV capacity estimation and real-time electricity market integration for DERs.

Asiedu, Shadrack (ORCID:0009000646004826)↗

Standardising the “Gregory method” for calculating equilibrium climate sensitivity

The equilibrium climate sensitivity (ECS) – the equilibrium global mean temperature response to a doubling of atmospheric CO 2 – is a high-profile metric for quantifying the Earth system's response to human-induced climate change. A widely applied approach to estimating the ECS is the “Gregory method” (Gregory et al., 2004), which uses an ordinary least squares (OLS) regression between the net radiative flux, N, and surface air temperature anomalies, ΔT, from a 150 year experiment in which atmospheric CO 2 concentrations are quadrupled. The ECS is determined by extrapolating the linear fit to N=0, i.e. the ΔT-intercept, indicating the point at which the system is back in equilibrium. This method has been used to compare ECS estimates across the CMIP5 and CMIP6 ensembles and will likely be a key diagnostic for CMIP7. Despite its widespread application, there is little consistency or transparency between studies in how the climate model data is processed prior to the regression, leading to potential discrepancies in ECS estimates. We identify 32 alternative data processing pathways, varying by differences in global mean weighting, net radiative flux variable, anomaly calculation method, and linear regression fit. Using 44 CMIP6 models, we systematically assess the impact of these choices on ECS estimates and calculate uncertainty ranges using two bootstrap approaches. While the inter-model ECS range is insensitive to the data processing pathway, individual outlier models exhibit notable differences. Approximating a model's native grid cell area (if irregular) with cosine of the latitude can decrease the ECS by 11 %, the choice of N-variable can change the ECS by 6 %, and some anomaly calculation methods can introduce spurious temporal correlations in the processed data. Beyond data processing choices, we also evaluate an alternative linear regression method – total least squares (TLS) – which has a more statistically robust basis than OLS. However, for consistency with previous literature, and given TLS may reduce the ECS compared to OLS (by up to 24 %), thereby making a known bias in the Gregory method worse, we do not feel there is sufficient clarity to recommend a transition to TLS in all cases. To improve reproducibility and comparability in future studies, we recommend a standardised Gregory method: weighting the global mean by cell area, using the top of the atmosphere (as opposed to the top of model) N-variable, and calculating anomalies by first applying a rolling average to the preindustrial control timeseries then subtracting from the raw CO 2 quadrupling experiment. This approach accounts for model drift while reducing noise in the data to best meet the pre-conditions of the linear regression. While CMIP6 results of the multi-model mean ECS appear insensitive to these processing choices, similar assumptions may not hold for CMIP7, underscoring the need for standardised data preparation in future climate sensitivity assessments.

Geosciences↗

Planetary Boundary-Layer Height (PBLHT) Value-Added Product: Remote-Sensing Retrievals

The planetary boundary layer (PBL) is fundamental to numerous atmospheric processes, including aerosol mixing and transport, cloud evolution, and precipitation formation. A critical parameter in these studies is the PBL height (PBLHT). This vertical depth is essential for characterizing PBL structures in numerical simulations and serves as a primary metric for estimating flux exchanges between the Earth’s surface and the atmosphere. Radiosonde (SONDE) observations provide high-vertical-resolution measurements of temperature and moisture profiles and are widely used to estimate PBLHT (Liu and Liang 2010, Seidel et al. 2010). The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility’s PBLHT value-added product (VAP) for radiosonde measurements, known as PBLHTSONDE, applies three commonly used methods—the Heffter (1980) method, the Liu and Liang (2010) method, and the bulk Richardson number approach (Seibert et al. 2000)—to derive PBLHT. The PBLHTSONDE VAP operates routinely at ARM observatories and mobile facilities, with data available from the ARM Data Center shortly after sounding observations are collected (Sivaraman et al. 2013). However, radiosonde observations are limited by their low temporal resolution. Most stations launch soundings only twice daily, which constrains the ability to investigate and characterize the temporal evolution of the PBL using radiosonde data alone. The use of continuous remote-sensing observations provides high temporal resolution of PBLHT estimates. These observations include aerosol lidars (Dang et al. 2019, Su et al. 2020), Doppler lidar (DL; Tucker et al. 2009, Krishnamurthy et al. 2021), and water vapor and/or temperature lidars and radiometers (Turner et al. 2014). These observations provide valuable data on the PBL’s thermodynamic properties (e.g., water vapor and/or temperature lidars and radiometers), dynamic properties (e.g., DL), and distribution of tracer substances (e.g., aerosol lidars), all of which can be used to estimate PBLHT. ARM developed PBLHT estimates from the micropulse lidar (MPL; PBLHTMPL), Doppler lidar (PBLHTDL), and combined Raman lidar (RL)/atmospheric emitted radiance interferometer (AERI) thermodynamic profiles (PBLHTTHERMO). Each estimate captures different physical characteristics of the boundary layer—aerosol tracers, vertical velocity turbulence, and thermodynamic structure—and exhibits distinct strengths and limitations depending on the PBL regime and time of day. In addition, the ARM ceilometer (CEIL) provides three potential PBLHT candidates derived from the vendor's built-in algorithm. Building on these individual retrievals, ARM developed the PBLHTBEML VAP, which combines the four remote-sensing-based estimates with ancillary meteorological variables using the machine learning approach of Zhang et al. (2025) to produce a best-estimate PBLHT at 10-minute resolution.

54 ENVIRONMENTAL SCIENCES↗

Architecture for Web-Based Visualization of Large-Scale Energy Domains: Preprint

With the growing penetration of inverter-based distributed energy resources and increased loads through electrification, power systems analyses are becoming more important and more complex. Moreover, these analyses increasingly involve the combination of interconnected energy domains with data that are spatially and temporally increasing in scale by orders of magnitude, surpassing the capabilities of many existing analysis and decision-support systems. We present the architectural design, development, and application of a high-resolution web-based visualization environment capable of cross-domain analysis of tens of millions of energy assets, focusing on scalability and performance. Our system supports the exploration, navigation, and analysis of large data from diverse domains such as electrical transmission and distribution systems, mobility and electric vehicle charging networks, communications networks, cyber assets, and other supporting infrastructure. We evaluate this system across multiple use cases, describing the capabilities and limitations of a web-based approach for high-resolution energy system visualizations.

grid modernization↗

DyG-DPCD: A Distributed Parallel Community Detection Algorithm for Large-Scale Dynamic Graphs

Dynamic (Temporal) graphs capture the valuable evolution of real-world systems, from the continuously evolving patterns of social interactions and genetic pathways to the dynamic fluctuations of economic forces. Detecting communities for such evolving networks poses unique challenges. Detecting and analyzing the evolution of communities within dynamic graphs unlocks valuable insights into the underlying structural and temporal patterns of real-world systems. However, the sheer volume of modern graph data and the inherent complexity of the temporal dimension pose significant challenges to scalable community detection algorithms. Addressing this gap, our work explores the limited landscape of scalable distributed-memory parallel methods specifically designed for dynamic network community detection. We propose a novel parallel algorithm, DyG-DPCD (Dynamic Graph Distributed Parallel Community Detection), to detect communities in dynamic networks using the Message Passing Interface (MPI) framework. We present a vertex-centric approach, allowing us to detect communities through local optimization. Furthermore, we enhance our baseline algorithm by incorporating three heuristics, which improve the algorithm’s performance significantly while maintaining the quality of the solutions. We demonstrate the efficiency of our algorithm by experimenting on several real-world large-scale networks with hundreds of millions of edges spanning diverse domains. Notably, DyG-DPCD achieves speedups between 25× and 30× for large networks that we experimented on using NERSC compute nodes. In conclusion, our algorithm outperforms the STINGER parallel re-agglomeration algorithm by 30×.

97 MATHEMATICS AND COMPUTING↗

Urban Land Surface Temperature Downscaling in Chicago: Addressing Ethnic Inequality and Gentrification

In this study, we developed a XGBoost-based algorithm to downscale 2 km-resolution land surface temperature (LST) data from the GOES satellite to a finer 70 m resolution, using ancillary variables including NDVI, NDBI, and DEM. This method demonstrated a superior performance over the conventional TsHARP technique, achieving a reduced RMSE of 1.90 °C, compared to 2.51 °C with TsHARP. Our approach utilizes the geostationary GOES satellite data alongside high-resolution ECOSTRESS data, enabling hourly LST downscaling to 70 m—a significant advancement over previous methodologies that typically measure LST only once daily. Applying these high-resolution LST data, we examined the hottest days in Chicago and their correlation with ethnic inequality. Our analysis indicated that Hispanic/Latino communities endure the highest LSTs, with a maximum LST that is 1.5 °C higher in blocks predominantly inhabited by Hispanic/Latino residents compared to those predominantly occupied by White residents. This study highlights the intersection of urban development, ethnic inequality, and environmental inequities, emphasizing the need for targeted urban planning to mitigate these disparities. The enhanced spatial and temporal resolution of our LST data provides deeper insights into diurnal temperature variations, crucial for understanding and addressing the urban heat distribution and its impact on vulnerable communities.

Lee, Jangho (ORCID:0000000289421092)↗

The temporal onset of associations of cortical proteins with cognitive resilience vary during late life

Background: Cortical proteins associated with cognitive resilience have been identified but their temporal onset in older adults is unknown. We present a multistage approach to first identify cortical proteins associated with cognitive resilience and then examine their associated temporal onset. Methods: We used data from a subset of 1088 decedents from two cohort-studies who had selected reaction monitoring proteomics from the dorsolateral prefrontal cortex, and at least 3 cognitive assessments. Cognition was assessed using a composite derived from 19 tests. We first used linear mixed-effects models to identify cortical proteins associated with cognitive resilience. We then used functional mixed-effects models to examine non-linear associations between proteins and cognitive resilience to identify their temporal onset. Results: Mean age at death was 90 years (SD = 6.4); 69 % were female. On average, cognition started to decline at around 15 years before death, with accelerated decline in the last 7 years. We identified 40 proteins associated with cognitive resilience, of which 17 proteins also showed non-linear associations. Non-linear associations indicated that higher levels of 10 proteins were associated with slower cognitive decline between 23 and 4 years before death. In contrast, higher levels of 7 proteins were associated with faster decline only within the last 7 years before death. Conclusions: Cognitive resilience proteins are differentially related to late-life cognitive aging; the onset of proteins that maintain cognition may begin many years before the onset of proteins that hasten cognitive decline. The temporal onset of cognitive resilience proteins may be crucial for timing efficacious interventions.

Zammit, Andrea↗

Temporal covariation of island arc Sr isotopes and seawater chemistry over the past 2 billion years

The chemical compositions of island arc basalts (IAB) reflect contributions from the mantle as well as fluids and melts from the subducting slab. Addition of radiogenic seawater Sr to oceanic crust through hydrothermal alteration and subsequent subduction is often invoked to explain elevated 87 Sr/ 86 Sr signatures in modern IAB. However, changes in the 87 Sr/ 86 Sr of island arc magmatic rocks through time has not been investigated, limiting our understanding of the factors influencing the Sr budgets of arcs throughout Earth’s history. To address this, we compiled 87 Sr/ 86 Sr values from island arc magmatic rocks ranging in age from modern to Paleoproterozoic, only including data from island arc localities that best preserve initial magmatic 87 Sr/ 86 Sr. Median initial 87 Sr/ 86 Sr values are consistently elevated compared to depleted mantle 87 Sr/ 86 Sr over this period, indicating persistent enrichment in radiogenic Sr in island arcs. Moreover, the elevation in island arc 87 Sr/ 86 Sr relative to the depleted mantle is variable. A notable rise in island arc 87 Sr/ 86 Sr during the late Neoproterozoic coincides with a steep increase in seawater 87 Sr/ 86 Sr and Sr concentration. To investigate this potential connectivity, we modeled the 87 Sr/ 86 Sr of island arc magmas between 0 and 830 Ma with inputs of depleted mantle 87 Sr/ 86 Sr, seawater 87 Sr/ 86 Sr, and seawater Sr concentration. The model reproduces the overall trajectory of the compiled data. We interpret the observed temporal variation in island arc 87 Sr/ 86 Sr values and its close association with fluctuations in seawater chemistry as evidence that changes in marine geochemistry have strongly influenced the Sr isotopic record of island arc magmas over time.

Science & Technology - Other Topics↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING↗