Search NASA⌕ Search

SEARCH · Search NASA

Results for “synthetic dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Validation of the DESI 2024 Lyα forest BAO analysis using synthetic datasets

The first year of data from the Dark Energy Spectroscopic Instrument (DESI) contains the largest set of Lyman-α (Lyα) forest spectra ever observed. This data, collected in the DESI Data Release 1 (DR1) sample, has been used to measure the Baryon Acoustic Oscillation (BAO) feature at redshift z = 2.33. In this work, we use a set of 150 synthetic realizations of DESI DR1 to validate the DESI 2024 Lyα forest BAO measurement presented in [1]. The synthetic data sets are based on Gaussian random fields using the log-normal approximation. We produce realistic synthetic DESI spectra that include all major contaminants affecting the Lyα forest. The synthetic data sets span a redshift range 1.8 < z < 3.8, and are analyzed using the same framework and pipeline used for the DESI 2024 Lyα forest BAO measurement. To measure BAO, we use both the Lyα auto-correlation and its cross-correlation with quasar positions. We use the mean of correlation functions from the set of DESI DR1 realizations to show that our model is able to recover unbiased measurements of the BAO position. We also fit each mock individually and study the population of BAO fits in order to validate BAO uncertainties and test our method for estimating the covariance matrix of the Lyα forest correlation functions. Finally, we discuss the implications of our results and identify the needs for the next generation of Lyα forest synthetic data sets, with the top priority being to simulate the effect of BAO broadening due to non-linear evolution.

79 ASTRONOMY AND ASTROPHYSICS↗

Validation of the DESI DR2 Ly⁢ 𝛼 BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-α (Ly α) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the baryonic acoustic oscillation (BAO) analysis using Ly α forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Lyα mocks that use a quasilinear input power spectrum to incorporate the nonlinear broadening of the BAO peak. Here, we have also increased the number of realizations used in the validation to 400, compared to the 150 realizations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask damped Lyman-α absorbers in our spectra. The BAO measurement from the Ly α dataset of DESI DR2 is presented in a companion publication.

Casas, L. [Institut de Física d’Altes Energies (IF↗

Validation of the DESI DR2 Ly$\alpha$ BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-$\alpha$ (Ly$\alpha$) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the Baryonic Acoustic Oscillation (BAO) analysis using Ly$\alpha$ forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Ly$\alpha$ mocks that use a quasi-linear input power spectrum to incorporate the non-linear broadening of the BAO peak. We have also increased the number of realisations used in the validation to 400, compared to the 150 realisations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask Damped Lyman-$\alpha$ Absorbers (DLAs) in our spectra. The BAO measurement from the Ly$\alpha$ dataset of DESI DR2 is presented in a companion publication.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Synthetic Streamflow Datasets to Support Emulation of Water Allocations via LSTM

This archive is the data companion to the bonney_et-al_2026_erc metarepo which generates synthetic data, trains an LSTM model, and generates performance metrics on the trained model. While the generation of the synthetic data is fully reprodicible, it is a computationally expensive process. This data archive contains the synthetic datasets needed for training and testing an LSTM model and reproduction of figures and tables. In addition, supplemenatary data products generating and visualizing results is also included, such as geospatial data for the basin. Contents There are two high level directories: `WRAP_archive/` and `repo_data/`. The `WRAP_archive` directory contains compressed intermediate dataproducts from the dataset generation workflow (marked as "I_Dataset_Generation" in the metarepo). These data products are not required by any scripts in the metarepo, but they are archived as they are expensive to generate and may have useful information for other analyses. The `repo_data` directory contains the necessary data for reproducing the workflow in the metarepo and should be decompressed and moved into the top level of the metarepo. Additional details are provided in README.md.

drought↗

Genesis Mission-Enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE)

Argonne National Laboratory is supporting the U.S. Department of Transportation’s (USDOT’s) Bureau of Transportation Statistics (BTS) with collaborative research on development and application of privacy preserving AI frameworks that leverage unmatched AI expertise and secure computing resources made available through the U.S. Genesis Mission1 . This research advances U.S. energy security goals by supporting a safe offshore energy industry with secure, domain-specific AI tools to analyze confidential industry datasets collected by BTS to rapidly improve identification of hazards, precursors, and systemic safety risks in high-risk operational environments. The staged, security-first approach begins with development and testing of Argonne’s Genesis Mission-enabled Secure AI to Fortify Energy Process Safety (Genesis-SAFE) framework within Argonne’s accredited secure computing enclave (ABLE) leveraging Argonne’s AI scientific assistant substrate (AISAC). Methods to build synthetic datasets were developed together with BTS for use in preparing synthetic datasets that can be used to validate data containment, governance, and security controls in the ABLE environment. Future research directions would focus on applying the Genesis-SAFE framework to CIPSEA-protected datasets entirely within ABLE to support confidentiality-preserving analysis of safety risks, trends, and contributing factors.

Kim, Hyekyung [Argonne National Laboratory (ANL), ↗

Synthetic spectra for Lyman- α forest analysis in the Dark Energy Spectroscopic Instrument

Synthetic data sets are used in cosmology to test analysis procedures, to verify that systematic errors are well understood and to demonstrate that measurements are unbiased. In this work we describe the methods used to generate synthetic datasets of Lyman-α quasar spectra aimed for studies with the Dark Energy Spectroscopic Instrument (DESI). In particular, we focus on demonstrating that our simulations reproduces important features of real samples, making them suitable to test the analysis methods to be used in DESI and to place limits on systematic effects on measurements of Baryon Acoustic Oscillations (BAO). We present a set of mocks that reproduce the statistical properties of the DESI early data set with good agreement. Additionally, we use a synthetic dataset to forecast the BAO scale constraining power of the completed DESI survey through the Lyman-α forest.

79 ASTRONOMY AND ASTROPHYSICS↗

Synthetic Streamflow Datasets Derived from DOE 9505 for Select Texas Basins

This dataset is generated using a Bayesian Hidden Markov Model trained on the DOE 9505 streamflow projection ensemble. A set of 21,000 streamflow realizations are generated for the Colorado, Sabine, and Trinity river basins in Texas. Separate models are trained either using the full 9505 ensemble or a subset based on three hyperparameters: bias correction, downscaling, and hydrological model.

drought↗

High-resolution fully-polarimetric synthetic aperture radar dataset

Fully-polarimetric synthetic aperture radar (PolSAR) data contain a rich body of elementary scattering physics information that is critically valuable for a broad range of applications and scientific purposes. However, there is a lack of available high-resolution (< 0.3048-m) data available for PolSAR phenomenology research. This article introduces a high-resolution PolSAR data set collected and provided by Sandia National Laboratories (SNL). The data sets were collected to support studying high-resolution scattering physics from different types of clutter and applications such as polarimetric-based terrain classification.

West, Roger Derek↗

Vehicular Re-Identification from Uncontrolled Multiple Views

Vehicle re-identification (re-ID) across disparate sensing modalities remains a fundamental challenge for transportation research. In this work, we introduce a deep multi-view vehicle re-ID framework that leverages Siamese networks to compare pairs of vehicle images and produce matching scores, enabling robust association across drastically different viewpoints such as those from UAVs, surveillance cameras, and ground sensors. The model exploits convolutional neural networks to learn features that remain discriminative under changes in angle, distance, and illumination, supporting more generalizable re-ID performance. As part of this effort, we also developed an automated pipeline to synchronize roadside and UAV video streams, producing a multi-perspective dataset that complements preexisting real collections and a synthetic dataset generated in this study. Together, these contributions advance the capability to re-identify vehicles across wide viewing baselines; establish a foundation for scalable, reproducible research in vehicle re-ID; and open pathways for future applications, such as inferring routine behaviors, movement patterns, and daily habits of the individual associated with the vehicle.

convolutional neural networks↗

Validation of the DESI 2024 Lyman alpha forest BAL masking strategy

Broad absorption line quasars (BALs) exhibit blueshifted absorption relative to a number of their prominent broad emission features. These absorption features can contribute to quasar redshift errors and add absorption to the Lyman-α (Lyα) forest that is unrelated to large-scale structure. We present a detailed analysis of the impact of BALs on the Baryon Acoustic Oscillation (BAO) results with the Lyα forest from the first year of data from the Dark Energy Spectroscopic Instrument (DESI). The baseline strategy for the first year analysis is to mask all pixels associated with all BAL absorption features that fall within the wavelength region used to measure the forest. We explore a range of alternate masking strategies and demonstrate that these changes have minimal impact on the BAO measurements with both DESI data and synthetic data. This includes when we mask the BAL features associated with emission lines outside of the forest region to minimize their contribution to redshift errors. We identify differences in the properties of BALs in the synthetic datasets relative to the observational data, as well as use the synthetic observations to characterize the completeness of the BAL identification algorithm, and demonstrate that incompleteness and differences in the BALs between real and synthetic data also do not impact the BAO results for the Lyα forest.

Lyman alpha forest↗

Resource-Adaptive Federated Text Generation with Differential Privacy

In cross-silo federated learning (FL), sensitive text datasets remain confined to local organizations due to privacy regulations, making repeated training for each downstream task both communication-intensive and privacy-demanding. A promising alternative is to generate differentially private (DP) synthetic datasets that approximate the global distribution and can be reused across tasks. However, pretrained large language models (LLMs) often fail under domain shift, and federated finetuning is hindered by computational heterogeneity: only resource-rich clients can update the model, while weaker clients are excluded, amplifying data skew and the adverse effects of DP noise. We propose a flexible participation framework that adapts to client capacities. Strong clients perform DP federated finetuning, while weak clients contribute through a lightweight DP voting mechanism that refines synthetic text. To ensure the synthetic data mirrors the global dataset, we apply control codes (e.g., labels, topics, metadata) that represent each client’s data proportions and constrain voting to semantically coherent subsets. This two-phase approach requires only a single round of communication for weak clients and integrates contributions from all participants. Experiments show that our framework improves distribution alignment and downstream robustness under DP and heterogeneity.

Wang, Jiayi [ORNL]↗

Alcock-Paczynski Blinding Scheme for the Ly-$α$ Forest Analysis

We present and validate a blinding method for the Lyman-$α$ (Ly$α$) forest analysis based on a modification of the Alcock-Paczynski test. In order to hide the background expansion history, the method employs a geometrical shift of each quasar (QSO) forest in wavelength space, once the quasar continuum has been fitted and the fluctuation field is extracted. The redshift positions for the QSO sample are also changed in a consistent manner. We show that the method remains effective when applied to real data, where contamination from metals and Lyman-$β$ is intrinsically mixed with the Lyman-$α$ forest. This limitation is primarily visible in the 1D correlation function, where other blinding strategies can mitigate the effect. To assess its effectiveness, the prescription is tested against a series of datasets of increasing complexity: from idealized low-noise mocks, to realistic DESI year one synthetic datasets, and finally to data from DESI first data release (DR1), using both the auto (Ly$α\times$Ly$α$) and cross (Ly$α\times$ QSO) correlations. We find that the method robustly shifts the BAO peak position from the 3D correlation functions to the expected value for cosmology changes of around 5% in the matter content, without altering the shape of the posteriors in the model parameters. In conclusion, this catalog-level blinding strategy is a viable method for cosmological inference with the Lyman-$α$ forest, particularly if a cross-analysis with other tracers using the same blinding strategy is pursued.

Perez-Sanchez, G. [Guanajuato U.] (ORCID:000900096↗

A novel approach to increase accuracy in remotely sensed evapotranspiration through basin water balance and flux tower constraints

Remote sensing-derived evapotranspiration (RSET) products capture the spatiotemporal variations of evapotranspiration (ET) from field to basin scales with unprecedented details. However, their accuracy varies across RSET estimation methods and diverse hydroclimate regions. While ET modeling efforts to account for biophysical processes and controlling parameters have made good progress in recent years, a parallel approach of integrating in-situ ET with RSET could reduce biases in RSET products. Basin water balance ET (WBET) and flux tower ET are widely applied to evaluate RSET accuracy, yet such ET measurements are rarely used for RSET bias corrections, especially for large area applications. To address this issue, we propose a novel approach: the water balance equivalence (WABE) method, which generates spatially continuous WBET for correcting biases in RSET products. The WABE method computes synthetic WBET by integrating observed WBET and flux tower-derived FLUXCOM ET, which fills the spatial gaps of observed WBET and generates a spatially continuous WBET dataset. Synthetic WBET (2002–2015 annual average) of eight-digit hydrologic unit code (HUC8) basins across the conterminous United States (CONUS), constituting 44 % (887 out of 2035 basins) of CONUS basins, was determined within 2.0 % (RMSE = 12 %) of observed WBET at CONUS and between 1–12 % (RMSE = 3–33 %) across 18 regions in CONUS. With WABE-based bias corrections, the overall annual bias of RSET decreased from 10 % (RMSE = 34 %) to 6 % (RMSE = 26 %) across 37 flux tower sites. The WABE method offers a new approach for RSET accuracy improvement and shows great promise for large area implementations with a potential to yield substantial benefits for building accurate basin water budgets and water management decisions.

Khand, Kul↗

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING↗

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY↗

DESI DR2 results. I. Baryon acoustic oscillations from the Lyman alpha forest

We present the baryon acoustic oscillation (BAO) measurements with the Lyman-𝛼 (Ly⁢𝛼) forest from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI) survey. Our BAO measurements include both the autocorrelation of the Ly⁢𝛼 forest absorption observed in the spectra of high-redshift quasars and the cross-correlation of the absorption with the quasar positions. The total sample size is approximately a factor of 2 larger than the DR1 dataset, with forest measurements in over 820,000 quasar spectra and the positions of over 1.2 million quasars. We describe several significant improvements to our analysis in this paper, and two supporting papers describe improvements to the synthetic datasets that we use for validation and how we identify damped Ly⁢𝛼 absorbers. Our main result is that we have measured the BAO scale with a statistical precision of 1.1% along and 1.3% transverse to the line of sight, for a combined precision of 0.65% on the isotropic BAO scale at 𝑧 eff =2.33. This excellent precision, combined with recent theoretical studies of the BAO shift due to nonlinear growth, motivated us to include a systematic error term in Ly⁢𝛼 BAO analysis for the first time. We measure the ratios 𝐷 𝐻 ⁡(𝑧 eff )/𝑟 𝑑 = 8.632 ± 0.098 ± 0.026 and 𝐷 𝑀 ⁡(𝑧 eff )/𝑟 𝑑 = 38.99 ± 0.52 ± 0.12, where 𝐷 𝐻 = 𝑐/𝐻⁡(𝑧) is the Hubble distance, 𝐷 𝑀 is the transverse comoving distance, 𝑟 𝑑 is the sound horizon at the drag epoch, and we quote both the statistical and the theoretical systematic uncertainty. The companion paper presents the BAO measurements at lower redshifts from the same dataset and the cosmological interpretation.

baryon acoustic oscillations↗

DESI DR2 Results I: Baryon Acoustic Oscillations from the Lyman Alpha Forest

We present the Baryon Acoustic Oscillation (BAO) measurements with the Lyman-alpha (LyA) forest from the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI) survey. Our BAO measurements include both the auto-correlation of the LyA forest absorption observed in the spectra of high-redshift quasars and the cross-correlation of the absorption with the quasar positions. The total sample size is approximately a factor of two larger than the DR1 dataset, with forest measurements in over 820,000 quasar spectra and the positions of over 1.2 million quasars. We describe several significant improvements to our analysis in this paper, and two supporting papers describe improvements to the synthetic datasets that we use for validation and how we identify damped LyA absorbers. Our main result is that we have measured the BAO scale with a statistical precision of 1.1% along and 1.3% transverse to the line of sight, for a combined precision of 0.65% on the isotropic BAO scale at $z_{eff} = 2.33$. This excellent precision, combined with recent theoretical studies of the BAO shift due to nonlinear growth, motivated us to include a systematic error term in LyA BAO analysis for the first time. We measure the ratios $D_H(z_{eff})/r_d = 8.632 \pm 0.098 \pm 0.026$ and $D_M(z_{eff})/r_d = 38.99 \pm 0.52 \pm 0.12$, where $D_H = c/H(z)$ is the Hubble distance, $D_M$ is the transverse comoving distance, $r_d$ is the sound horizon at the drag epoch, and we quote both the statistical and the theoretical systematic uncertainty. The companion paper presents the BAO measurements at lower redshifts from the same dataset and the cosmological interpretation.

79 ASTRONOMY AND ASTROPHYSICS↗

MIC-DP: A Scalable Correlation-Aware Differential Privacy Framework for High-Dimensional Data

Conventional differential privacy (DP) assumes record independence, limiting effectiveness on real-world datasets with temporal, spatial, or structural correlations. These dependencies undermine privacy guarantees and degrade utility in domains like healthcare, IoT, and smart city analytics. We propose Maximum Information Correlated Differential Privacy (MIC-DP), a novel framework that dynamically calibrates noise based on statistical dependencies. MIC-DP uses the Maximum Information Coefficient (MIC) to capture both linear and nonlinear correlations without explicit modeling, enabling adaptive sensitivity adjustment and improved privacy–utility trade-offs. Evaluations on healthcare (MIMIC), demographic (ACI), and synthetic datasets show that MIC-DP reduces mean absolute error (MAE) by up to 5.2% under strict privacy budgets (ϵ≤1), with aggregate utility improvements reaching 18% across datasets and evaluation metrics. MIC-DP provides formal (ϵ,δ)-privacy guarantees, scales efficiently with feature count, and supports deployment in moderate-scale, privacy-sensitive applications. Its tunable performance and runtime efficiency make MIC-DP suitable for privacy-sensitive applications where low-latency analytics and strong privacy guarantees must coexist. These results demonstrate MIC-DP’s effectiveness as a correlation-aware solution for practical DP.

Yang, Wenjun [Univ. of Washington, Tacoma, WA (Uni↗