Search NASA⌕ Search

SEARCH · Search NASA

Results for “data distributions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Enhancing Biopreparedness through a Model System to Understand the Molecular Mechanisms that Lead to Pathogenesis and Disease Transmission: NW-BRaVE

The science of biopreparedness to counter biological threats hinges on understanding the fundamental principles and molecular mechanisms that lead to pathogenesis and disease transmission. Our vision to address this challenge is to create a powerful and user-friendly platform to elucidate the fundamental principles of how molecular interactions drive pathogen-host relationships and host shifts. We will enable groundbreaking discoveries by integrating a wide range of structural, genomics, proteomics, and other advanced omics measurements, along with evolutionary and artificial intelligence predictions. To make sure the system is applicable to real-world problems, we will develop it in the context of a tractable model system, the small, abundant, and accessible photosynthetic cyanobacteria and their constantly co-adapting viral pathogens, cyanophages. This model will maintain the system’s applicability to real-world problems and techniques, but the overall focus will be on elucidating general principles of detecting, assessing, and surveilling molecular interaction, adaptation, and coevolution that are system agnostic and therefore extensible to other viral-host interactions. Our overall objectives are to (1) identify the molecular complexes that comprise the cyanobacteria redox macromolecular subsystem and how they dynamically change with bacteriophage infection in situ, using cryo-electron tomography; (2) profile regulatory changes during infection using proteomics, multiomics, and experimental validation, and integrate the data with in situ structures; (3) use genomics and metagenomics to determine environmental and population factors across time scales that impact the interactions between marine cyanobacteria and their cyanophage parasites, predicting the evolutionary origins of in situ structural and functional interactions, convergence and coevolution; and (4) develop a data integration and transformation platform that facilitates the integration of in situ, proteomic, and evolutionary measurements of molecular interactions to surveil diverse hosts and parasites in various environmental contexts. These objectives address Focus Area 2 Reveal Molecular Interactions Across Biological Scales for Design of Targeted Interventions. Our powerful and user-friendly platform will enhance connections between the often-siloed fields of structure, molecular phenotype, and evolutionary genomics that are key to biopreparedness, but in need of integration (Figure 1). We will build an integrated navigation tool to facilitate the effective use of globally distributed experimental data for integrated analysis and predictive modeling. The project will develop, implement, and test a platform to assess host-pathogen molecular interactions, adaptation to hosts and host shifts, and coevolution between hosts and pathogens, successfully impacting the research community by revolutionizing abilities to study any host-pathogen interaction, encourage diverse community contributions, and gain fundamental insights into how proteins adapt to new contexts. This ability will be critical for designing early interventions to address future threats. We will build surveillance training capability, aiming for a fair and equitable response to future pandemics and biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

Navigating Integration: Key Challenges for Data Centers, Nuclear Stakeholders, and Utility Operators

he exponential growth of data centers—driven by artificial intelligence and cloud computing—is reshaping the U.S. energy landscape, presenting urgent challenges and transformative opportunities for data center developers, nuclear energy providers, and utility operators. As data centers are projected to consume up to 12% of U.S. electricity by 2028, stakeholders must address rapid deployment needs, grid congestion, and the demand for reliable, high-quality power. This presentation explores the multifaceted barriers to integrating data centers with nuclear and utility infrastructure, including land use constraints, public perception, regulatory complexity, and workforce alignment. It highlights the distinct priorities and operational cultures of each sector, and the friction that arises from misaligned planning horizons and risk tolerances. We examine collaborative strategies such as co-siting, hybrid power-purchase agreements, unified community engagement, and innovative financing models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Fourier-MIONet: Fourier-enhanced multiple-input neural operators for multiphase modeling of geological carbon sequestration

Geologic carbon sequestration (GCS) is a safety-critical technology that aims to reduce the amount of carbon dioxide in the atmosphere, which also places high demands on reliability. Multiphase flow in porous media is essential to understand CO 2 migration and pressure fields in the subsurface associated with GCS. However, numerical simulation for such problems in 4D is computationally challenging and expensive, due to the multiphysics and multiscale nature of the highly nonlinear governing partial differential equations (PDEs). It prevents us from considering multiple subsurface scenarios and conducting real-time optimization. Here, we develop a Fourier-enhanced multiple-input neural operator (Fourier-MIONet) to learn the solution operator of the problem of multiphase flow in porous media. Fourier-MIONet utilizes the recently developed framework of the multiple-input deep neural operators (MIONet) and incorporates the Fourier neural operator (FNO) in the network architecture. Once Fourier-MIONet is trained, it can predict the evolution of saturation and pressure of the multiphase flow under various reservoir conditions, such as permeability and porosity heterogeneity, anisotropy, injection configurations, and multiphase flow properties. Compared to the enhanced FNO (U-FNO), the proposed Fourier-MIONet has 90% fewer unknown parameters, and it can be trained in significantly less time (about 3.5 times faster) with much lower CPU memory (<15%) and GPU memory (<35%) requirements, to achieve similar prediction accuracy. In addition to the lower computational cost, Fourier-MIONet can be trained with only 6 snapshots of time to predict the PDE solutions for 30 years. Furthermore, we observed that Fourier-MIONet can maintain good accuracy when predicting out-of-distribution (OOD) data. The excellent generalizability of Fourier-MIONet is enabled by its adherence to the physical principle that the solution to a PDE is continuous over time. Furthermore, the developed Fourier-MIONet makes it possible to solve the long-time evolution of geological carbon sequestration in a large-scale three-dimensional space accurately and efficiently.

97 MATHEMATICS AND COMPUTING↗

Three-Dimensional Grid Visualization for Planning Activities: A Dubai Case Study

National Laboratory of the Rockies (NLR), in collaboration with the Dubai Electricity and Water Authority (DEWA) and Infra-X, has undertaken the Energy Visualization Analysis Project. The aim of this project is to enhance analytical and 3D visualization capabilities for distribution network planning and renewable energy integration. As modern grid continues to evolve with large-scale solar PV deployment and emerging distributed energy resources (DERs), the ability to effectively analyze, visualize, and communicate complex grid behaviors has become increasingly critical. The project focuses on developing empirical use cases based on real distribution feeder data and engineering workflows, ensuring the outcomes are directly aligned with operational environment. Through time-series power flow simulations and nodal hosting capacity analysis, the study quantifies the impacts of high PV penetration on voltage and thermal limits within representative 11 kV feeders. These analyses identify specific nodes and conditions where DER integration challenges arise. Furthermore, a Battery Energy Storage System (BESS) optimization algorithm was applied to determine the optimal size and placement of storage systems that can mitigate network constraints and enhance hosting capacity. The comparative results between base-case and BESS-augmented scenarios clearly demonstrate improvements in network stability and load management efficiency. In parallel, the NLR team developed an immersive 3D visualization framework, enabling interactive exploration of grid simulations using commodity head-mounted display (HMD) systems. This framework transforms conventional 2D simulation data into spatially intuitive visual environments - allowing engineers to analyze feeder conditions, PV hosting potential, and BESS effects in real time. This report represents the first foundational phase in establishing a visualization-driven analytical ecosystem. It provides a methodological foundation for data integration, visualization architecture, and simulation-based decision support, paving the way for large-scale adoption of immersive visualization across DEWA's Smart Grid Initiative, R&D activities, and future network resilience studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

WHONDRS River Corridor Surface Water Metabolites and Geochemistry from Global Sites

This dataset supports a broader study examining the character of organic matter that may be delivered to subsurface sediments via hydrologic exchange. To implement the global survey, free stream sampling kits were provided to interested volunteers throughout the world. Samples were collected with minimal constraints in terms of location, but following strict protocols, and shipped for metabolomic analysis via Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). In addition, basic geochemistry analyses (e.g., dissolved organic matter concentration) were conducted, standardized photos of each field system were taken, and extensive metadata were captured. Sampling began in 2018 and is ongoing as of 2025. This dataset is comprised of one folders of field photos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; and (7) a subfolder with sample data. The sample data subfolder contains (1) surface water dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) methods codes; (3) surface water FTICR methods; and (4) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains three subfolders, one containing the.xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, or .png. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

Biogeochemistry↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

A Framework for Streamlining Large Load Interconnection: An Assessment of Current Practices and Emerging Strategies for Fair and Efficient Load Interconnection Processes

This report focuses on current perceptions, challenges, and processes related to interconnection of LELs to the grid; while these loads include several types of entities, the primary focus of this work is on data centers, the largest current driver of this growth. It begins by summarizing relevant context related to load growth in the United States, including key stakeholder perceptions as data center deployment drive historic load growth nationwide. It then reviews the current landscape of large load interconnection processes and challenges and then proposes a framework of potential strategies intended to support a more consistent, streamlined, and fair approach to load interconnection. This framework is adapted from solutions proposed or adopted to address similar challenges faced in generator interconnection queues as well as notable reforms currently being proposed or enacted for large loads that could be more widely adopted.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNN

We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier.

97 MATHEMATICS AND COMPUTING↗

High-temperature anomalous Hall effect driven by frustrated spin fluctuations in the antiferromagnetic delafossite metal PdCrO 2

Metallic delafossite oxides are of exceptional interest due to their ultraclean metallic transport. In the case of PdCrO 2 , this arises in a triangular-lattice antiferromagnet, creating a unique opportunity to study frustrated magnetism in a very clean metal. Here, we combine a chemical vapor transport crystal growth approach with magnetic, thermodynamic, magnetotransport, and neutron scattering measurements to elucidate the striking anomalous Hall effect (AHE) in antiferromagnetic PdCrO 2 . The unconventional AHE (with anomalous Hall conductivity ~10 5 Ω −1 cm −1 ) and a large positive magnetoresistance effect (>1,000%) are shown to exhibit complex temperature dependencies, persisting to almost seven times the Néel temperature (~250 K). These effects are directly compared to elastic neutron scattering, inelastic neutron scattering, and neutron magnetic pair distribution function data, establishing unambiguous links between anomalous magnetotransport properties and directly probed short-range spin fluctuations. The latter occur over a notably broad temperature range due to geometrical magnetic frustration. Connecting to recent experimental and theoretical developments, these findings are discussed in terms of a temperature-dependent interplay between chiral spin order and chiral spin fluctuations, significantly elucidating the high-temperature anomalous magnetotransport in such compounds.

anomalous Hall effect↗

Uncertainty Quantification for Neutron Shield Using Convolutional Neural Networks

Uncertainty quantification from radiation transport calculations was conducted using a Bayesian inference approach. A surrogate model, using a convolutional neural network, was employed to emulate the neutron fluence, which was simulated with a Monte Carlo radiation transport model. This allowed for a computationally cheap approach to evaluate input parameters and to sample their corresponding posterior probability distributions. Experimental data from the literature were employed to perform uncertainty quantification studies for concrete shields. As a result, the method is a nonintrusive approach that enables studies with multiple input parameters and can be applied to any radiation transport model.

Bayesian inference↗

Evidence of distinct nonpolar/polar ordering at long/short ranges in relaxor ferroelectrics

Distinct atomic nonpolar/polar ordering is determined at long, intermediate, and short ranges using synchrotron x-ray diffraction, Raman scattering, and pair distribution function data in K 0.5 Na 0.5 NbO 3 -doped BaTiO 3 -based relaxor ferroelectrics, viz., (1 - x) Ba 0.9 Sr 0.1 TiO 3 -x KNN 50. Two different polar phases with distinct symmetries are identified at short/intermediate ranges for x = 0.10, where a monoclinic phase at short ranges gradually transforms into a rhombohedral phase at intermediate ranges. Here, in contrast, a nonpolar phase with average cubic symmetry is found to be stable at long ranges for a wide temperature range (100 ≤ T ≤ 500 K). The polar behavior of short/intermediate-range ordering is quantified in terms of the amplitude of ferroelectric phonon mode $Γ^-_4$ (q = 0, 0, 0) (which corresponds to the zone center of the cubic Brillouin zone). The amplitude of the $Γ^-_4$ phonon mode increases with decreasing temperature, suggesting the enhancement of ferroelectric polarization at short/intermediate ranges on lowering temperatures. Enhanced ferroelectric polarization at low temperature is also visible and evident by ferroelectrostriction and is quantified by the increase in magnitude of spontaneous volume ferroelectrostriction. Consequently, Zero thermal expansion is observed at low temperatures.

36 MATERIALS SCIENCE↗

CAFE AU LAIT: Compute-Aware Federated Augmented Low-Rank AI Training

Federated finetuning is crucial for unlocking the knowledge embedded in pretrained Large Language Models (LLMs) when data are geographically distributed across clients. Unlike finetuning with data from a single institution, federated finetuning allows collaboration across multiple institutions, enabling the utilization of diverse and decentralized datasets while preserving data privacy. Given the high computing costs of LLM training and the emphasis on energy efficiency in Federated Learning (FL), Low-Rank Adaptation (LoRA) has emerged as a widely adopted algorithm due to its significantly reduced number of trainable parameters. However, this assumes that all data silos have the necessary computing resources to compute local updates of LLMs. Nevertheless, in practice, the computing resources across clients are highly heterogeneous: while some may have access to hundreds of GPUs, others might have limited or no GPU access. Recently, federated finetuning using synthetic data has been proposed, allowing clients to participate in a collaborative training run without training LLMs locally. However, our experimental results reveal a performance gap between models trained using synthetic data and those trained using local updates. Motivated by the observed heterogeneity in computing resources and the performance gap, we propose a novel two-stage algorithm that leverages the storage and computing capabilities of a strong server. In the first stage, under the coordination of the strong server, clients with limited computing resources collaborate to generate synthetic data, which is transferred to and stored on the strong server. In the second stage, the strong server uses this synthetic data on behalf of the resource-constrained clients to perform federated LoRA finetuning alongside clients with sufficient computing resources. This approach ensures that all clients can participate in the finetuning process. Experimental results demonstrate that incorporating local updates from even a small fraction of clients improves performance compared to using synthetic data for all clients. Furthermore, we incorporate the Gaussian mechanism in both stages to guarantee client-level differential privacy.

Wang, Jiayi [ORNL]↗

Dynamic Temporal Graph Sequence Data for Resilience-Oriented Distribution Network Reconfiguration

This dataset comprises temporal dynamic graph sequences generated from power grid simulations focused on grid reconfiguration to enhance resilience. The simulations model failure propagation under varying conditions, with nodes assigned distinct failure probabilities. For each time step, the dataset captures the evolution of node states (functional or failed) and features critical to grid operations, such as pv_output, load_profile, load_dispatch, dg_output, loss, and voltage. Node types include sources, normal loads, and nodes with specific equipment like PVs, micro turbines, or shunt capacitors. The dataset is structured to support the training of dynamic graph neural networks, facilitating research on node feature prediction and edge dynamics under failure scenarios. Three distinct configurations are included, providing a robust foundation for modeling power grid resilience.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Review of Grey Box/Black Box Data Contamination Metrics on Open and Commercial Models

Dataset contamination is a problem where benchmarks and tasks used to evaluate the capabilities of Large Language Models (LLMs) have been incorporated into the training dataset of the models. This gives a false sense of performance that can overestimate how these models will function on truly unseen data. This problem becomes worse with commercial LLMs with larger and non-accessible training data, so techniques have been developed to try to measure the degree to which a model is contaminated with a benchmark’s data. To understand the effectiveness of these techniques, particularly when evaluating contamination on coding tasks, we review trends and categorize techniques by the degree of access to the model that is required. The research literature on this topic has reported mixed effectiveness of these techniques, so we select a set of black box (text access only) and grey box (access to model loss/probabilities required) techniques and apply them to both commercial and non-commercial models. We implement these metrics as part of a framework to test the contamination of Python code in LLMs to see to what extent we can replicate the effectiveness (or ineffectiveness) of these contamination detection techniques. Though we find mixed results in the capabilities of these metrics to identify contamination, we do observe evidence that they can identify contamination (broadly) in fine-tuned models when both a baseline and fine-tuned model is present. Additionally, similarity metrics were able to identify between contaminated and uncontaminated data even in situations where the data is distributionally similar (e.g., drawn from the same set of code projects).

97 MATHEMATICS AND COMPUTING↗

Large Load Management and Grid Planning

This presentation will cover these topics, followed by a discussion and Q&A: Grid impacts of load growth of the current bulk power system, a review of recent trends and events related to large load integration, what might these new large loads look like, what are options for reducing costs and timelines of integration, and working toward a framework for identifying hot spots for efficiently interconnecting large loads.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Utah FORGE: Neubrex Well 16B(78)-32 DAS Data - April, 2024

This dataset comprises Distributed Acoustic Sensing (DAS) data collected from the Utah FORGE monitoring well 16B(78)-32 (the producer well) during hydraulic fracture stimulation operations conducted in April 2024. The data were acquired continuously over the stimulation period at a temporal sampling rate of 10,000 Hz (10 kS/s) and a spatial resolution of approximately 3.35 feet (1.02109 meters). The measurements were captured using a Neubrex NBX-S4100 Time Gated Digital DAS interrogator unit connected to a single-mode fiber optic cable, which was permanently installed within the casing string. All recorded channels correspond to downhole segments of the fiber optic cable, from a measured depth (MD) of 5,369.35 feet to 10,352.11 feet. The DAS data reflect raw acoustic energy generated by physical processes within and surrounding the well during stimulation activities at wells 16A(78)-32 and 16B(78)-32. These data have potential applications in analyzing cross-well strain, far-field strain rates (including microseismic activity), induced seismicity, and seismic imaging. Metadata embedded in the attributes of the HDF5 files include detailed information on the measured depths of the channels, interrogation parameters, and other acquisition details. The dataset also includes a recording of a seminar held on September 19, 2024, where Neubrex's Chief Operating Officer presented insights into the data collection, analysis, and preliminary findings. The raw data files, stored in HDF5 format, are organized chronologically according to the recording intervals from April 9 to April 24, 2024, with each file corresponding to a 12-second recording interval.

15 GEOTHERMAL ENERGY↗

A Data-Driven Voltage Control Strategy for Distribution Grids With Distributed Energy Resources

Traditionally, distribution system control approaches have been model-based. The deployment of advanced metering infrastructure has provided electric utilities with the capability of data-driven control with real-time measurements. The shift from model-based to data-driven control represents a significant advancement in the management of distribution systems, offering a more adaptive approach to system control because of the ability to dynamically adapt to changing conditions without the need for system modeling. Here, in this paper, a behavioral data-driven control method is developed to provide voltage regulation to an actual distribution system by controlling the legacy devices and distributed energy resource (DER) assets. The studied distribution system has a load tap changer and three capacitor banks as the legacy devices and photovoltaic systems as the DERs. The performance of the proposed control algorithm is validated using a laboratory test bed setup considering multiple scenarios. The results show that the proposed control achieved 99% voltage regulation.

24 POWER TRANSMISSION AND DISTRIBUTION↗