Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

From RNNs to Foundation Models: An Empirical Study on Commercial Building Energy Consumption

Accurate short-term energy consumption forecasting for commercial buildings is crucial for smart grid operations. While smart meters and deep learning models enable forecasting using past data from multiple buildings, data heterogeneity from diverse buildings can reduce model performance. The impact of increasing dataset heterogeneity in time series forecasting, while keeping size and model constant, is understudied. We tackle this issue using the ComStock dataset, which provides synthetic energy consumption data for U.S. commercial buildings. Two curated subsets, identical in size and region but differing in building type diversity, are used to assess the performance of various time series forecasting models, including finetuned open-source foundation models (FMs). The results show that dataset heterogeneity and model architecture have a greater impact on post-training forecasting performance than the parameter count. Moreover, despite the higher computational cost, finetuned FMs demonstrate competitive performance compared to base models trained from scratch.

commercial buildings↗

A platform to measure isentropes from proton-heated warm dense matter on short pulse laser facilities

We describe the development of an experimental platform that measures the release isentrope of materials heated isochorically to temperatures of a few electron volts, using short-pulse laser-produced protons to heat the sample and long-pulse laser-produced x rays to perform streaked x-ray radiography. The density profiles derived from the radiography data are integrated to generate pressure–density isentropes, independent of prior knowledge of the equation of state of the sample material. In order to understand the sensitivities of isentrope extraction from radiography data, we analyze synthetic radiographs generated by a radiation hydrodynamics code. Noise reduction and high spatial resolution are critical for isentrope reconstruction, as demonstrated by the analysis of a proof-of-principle shot day on the OMEGA-EP facility. In conclusion, the data demonstrate the feasibility of the platform for characterizing isentropes, and we discuss the necessary improvements to enhance precision in differentiating between equation-of-state models.

Equations of state↗

Validation of the DESI DR2 Ly⁢ 𝛼 BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-α (Ly α) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the baryonic acoustic oscillation (BAO) analysis using Ly α forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Lyα mocks that use a quasilinear input power spectrum to incorporate the nonlinear broadening of the BAO peak. Here, we have also increased the number of realizations used in the validation to 400, compared to the 150 realizations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask damped Lyman-α absorbers in our spectra. The BAO measurement from the Ly α dataset of DESI DR2 is presented in a companion publication.

Casas, L. [Institut de Física d’Altes Energies (IF↗

Validation of the DESI DR2 Ly$\alpha$ BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-$\alpha$ (Ly$\alpha$) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the Baryonic Acoustic Oscillation (BAO) analysis using Ly$\alpha$ forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Ly$\alpha$ mocks that use a quasi-linear input power spectrum to incorporate the non-linear broadening of the BAO peak. We have also increased the number of realisations used in the validation to 400, compared to the 150 realisations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask Damped Lyman-$\alpha$ Absorbers (DLAs) in our spectra. The BAO measurement from the Ly$\alpha$ dataset of DESI DR2 is presented in a companion publication.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A DATA EFFICIENT SPARSE MODELING FRAMEWORK FOR POWER ESTIMATION IN WATER TREATMENT SENSING OPERATIONS

With increasing freshwater scarcity, advanced process design mechanisms such as Closed-Circuit Reverse Osmosis (CCRO) and Digital/Physical Twin systems are gaining traction in water treatment and reuse operations. While digital and physical twin models enable improved system insight and control, their development is often expensive and computationally intensive, requiring large volumes of synthetic or experimental data to characterize underlying process dynamics. This work introduces a sparse surrogate modeling framework to estimate power consumption from measured flow and pressure variables, along with their nonlinear polynomial and interaction expansions. To ensure model reliability and reduce overfitting, a two-stage pipeline is proposed. First, a dynamic data filtering algorithm is employed to remove uninformative observations and transient operational states. Second, a sparse penalized regression technique is applied to select a minimal set of parsimonious features. The proposed model achieves high sparsity, retaining only 7 out of 34 candidate features (≈79.41% sparsity) while delivering a root mean square error (RMSE) of 0.072 on the test dataset.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Development of Machine Learning Algorithm for Pebble Bed Modular Reactor Misuse Detection

The objective of this work was to develop a machine learning ensemble that could assist pebble bed reactor verification by evaluating whether a given pebble circulating through a PBR was normal or anomalous using gamma spectroscopy measurements from a notional PBR burnup measurement system. Using a PBR reference design, data sets of synthetic gamma spectra representative of BUMS measurements of normal and anomalous pebbles that may be used to produce special fissile material were generated to train and test an ML anomaly detection ensemble on two reference scenarios – substitution of normal pebbles with target pebbles for production of Pu or 233 U. The ML ensemble correctly identified all anomalous pebbles in the testing data set, and while perfect ensemble performance is normally indicative of overfitting, it was concluded that significantly lower photon intensity of target pebbles produced distinctly less intense photon spectra to where perfect ensemble performance was expected.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Continuous-variable quantum Boltzmann machine

Here, we propose a continuous-variable quantum Boltzmann machine (CVQBM) using a powerful energy-based neural network. It can be realized experimentally on a continuous-variable (CV) photonic quantum computer. We used a CV quantum imaginary time evolution (QITE) algorithm to prepare the essential thermal state and then designed the CVQBM to proficiently generate continuous probability distributions. We applied our method to both classical and quantum data. Using real-world classical data, such as synthetic-aperture radar (SAR) images, we generated probability distributions. For quantum data, we used the output of CV quantum circuits. We obtained high fidelity and low Kullback–Leibler (KL) divergence showing that our CVQBM learns distributions from given data well and generates data sampling from that distribution efficiently. We also discussed the experimental feasibility of our proposed CVQBM. Our method can be applied to a wide range of real-world problems by choosing an appropriate target distribution (corresponding to, e.g., SAR images, medical images, and risk management in finance). Moreover, our CVQBM is versatile and could be programmed to perform tasks beyond generation, such as anomaly detection.

SAR images↗

Bayesian event categorization matrix approach for explosion monitoring

Current efforts to correctly categorize natural events from suspected explosion sources with data that is collected by ground- or space-based sensors presents historical challenges that remain unaddressed by the Event Categorization Matrix (ECM) model. Smaller historical events (lower yield explosions) may have data available from fewer measurement techniques than are available today, and therefore, a historical event record can lack a complete set of discriminants. The covariance structures can also differ between such observations of event (source-type) categories. Both obstacles are problematic for the classic ECM model. Our work addresses this gap and presents a Bayesian update to the previous ECM model, termed the Bayesian Event Categorization Matrix model, which can be trained on partial observations and does not rely on a pooled covariance structure. We further augment the ECM model with Bayesian Decision Theory so that false negative or false positive rates of an event categorization can be reduced in an intuitive manner. To demonstrate improved categorization rates for the Bayesian Event Categorization Matrix model, we compare an array of Bayesian and classic models with multiple performance metrics using Monte Carlo experiments. We use both synthetic and real data. Our Bayesian models show consistent gains in overall accuracy and lower false negative rates relative to the classic ECM model. Here, we propose future avenues to improve Bayesian Event Categorization Matrix models’ decision making and predictive capability.

58 GEOSCIENCES↗

Identifying microbial drivers in biological phenotypes with a Bayesian network regression model

Abstract In Bayesian Network Regression models, networks are considered the predictors of continuous responses. These models have been successfully used in brain research to identify regions in the brain that are associated with specific human traits, yet their potential to elucidate microbial drivers in biological phenotypes for microbiome research remains unknown. In particular, microbial networks are challenging due to their high dimension and high sparsity compared to brain networks. Furthermore, unlike in brain connectome research, in microbiome research, it is usually expected that the presence of microbes has an effect on the response (main effects), not just the interactions. Here, we develop the first thorough investigation of whether Bayesian Network Regression models are suitable for microbial datasets on a variety of synthetic and real data under diverse biological scenarios. We test whether the Bayesian Network Regression model that accounts only for interaction effects (edges in the network) is able to identify key drivers (microbes) in phenotypic variability. We show that this model is indeed able to identify influential nodes and edges in the microbial networks that drive changes in the phenotype for most biological settings, but we also identify scenarios where this method performs poorly which allows us to provide practical advice for domain scientists aiming to apply these tools to their datasets. BNR models provide a framework for microbiome researchers to identify connections between microbes and measured phenotypes. We allow the use of this statistical model by providing an easy‐to‐use implementation which is publicly available Julia package at https://github.com/solislemuslab/BayesianNetworkRegression.jl .

59 BASIC BIOLOGICAL SCIENCES↗

Interpolation of computed gamma-ray detector response functions

Gamma-ray spectra measured by traditional detectors contain features that result from a combination of the effects of detector materials/geometry, the incident gamma-ray energy, and the angle of entry. The features, such as the full-energy photopeak, Compton continuum, annihilation peak, and escape peaks, are governed by simple relationships depending on incident energy and have been known for a long time. Monte Carlo computer simulations of gamma rays interacting with a detector will show these features, and with a resolution function applied, the results should look similar to real measurements. The traditional approach to creating a detector response function requires many separate simulations of monoenergetic gamma rays striking the detector. This paper presents a new approach to developing computed detector response functions. The new approach involves a much smaller number of monoenergetic gamma-ray simulations and uses interpolation to quickly generate the responses of gamma rays that were not simulated. During the interpolation process, the underlying physics equations are used to accurately compute the response of a given energy gamma ray from the small set of simulations. Such work enables accelerated generation of synthetic radiation detector data.

Detector response↗

A physics-constrained deep learning surrogate model of the runaway electron avalanche growth rate

A surrogate model of the runaway electron avalanche growth rate in a magnetic fusion plasma is developed. This is accomplished by employing a physics-informed neural network (PINN) to learn the parametric solution of the adjoint to the relativistic Fokker–Planck equation. The resulting PINN is able to evaluate the runaway probability function across a broad range of parameters in the absence of any synthetic or experimental data. This surrogate of the adjoint relativistic Fokker–Planck equation is then used to infer the avalanche growth rate as a function of the electric field, synchrotron radiation and effective charge. Predictions of the avalanche PINN are compared against first principle calculations of the avalanche growth rate with excellent agreement observed across a broad range of parameters.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Life Cycle Cost Analysis of Prestressed Concrete Poles Subjected to Wind, Surges, and Waves

Prestressed concrete (PC) poles are becoming popular choices to support coastal power transmission systems. However, the existing literature does not offer a detailed analysis of the effectiveness of PC poles in terms of long-term vulnerabilities and the direct and indirect costs. This is due to (1) lack of fragility models for PC poles and (2) lack of probabilistic wind, storm surge, and wave models in coastal settings. In this study, we address these gaps through a series of Monte Carlo simulations to estimate fragility of PC poles as a function of age and hazard (wind, surges, and waves) intensity, and the development of a probabilistic hazard model based on 10,000 years of synthetic tropical cyclone data. The probabilistic hazard model is used in conjunction with high-resolution hydrodynamic models to generate realizations of coastal wind, storm surges, and waves for the Louisiana and Mississippi coasts. A comprehensive life cycle cost analysis for a service life of 70 years considering direct and indirect losses is conducted to compare the performance of a transmission line located in Pascagoula, Mississippi, when wood poles are replaced by PC poles. Results showed that aging has a minor effect on the reliability of PC poles, highlighting the advantages of replacing wood poles with PC poles, especially in coastal areas. In addition, PC poles are significantly more cost-effective compared with wood poles over their life cycle, leading to an estimated saving of $11.55 million (68.17% reduction). The results of this study provide key insight to inform decision-making processes to keep the coastal power grids resilient and cost-effective against future storm hazards.

natural disasters↗

A physics-informed deep learning description of Knudsen layer reactivity reduction

A physics-informed neural network (PINN) is used to evaluate the fast ion distribution in the hot spot of an inertial confinement fusion target. The use of tailored input and output layers to the neural network is shown to enable a PINN to learn the parametric solution to the Vlasov–Fokker–Planck equation in the absence of any synthetic or experimental data. As an explicit demonstration of the approach, the specific problem of Knudsen layer fusion yield reduction is treated. Here, the predictions from the Vlasov–Fokker–Planck PINN are used to provide a non-perturbative solution of the fast ion tail in the vicinity of the hot spot, thus allowing the spatial profile of the fusion reactivity to be evaluated for a range of collisionalities and hot spot conditions. Excellent agreement is found between the predictions of the Vlasov–Fokker–Planck PINN and the results from traditional numerical solvers with respect to both the energy and spatial distribution of fast ions and the fusion reactivity profile, demonstrating that the Vlasov–Fokker–Planck PINN provides an accurate and efficient means of determining the impact of Knudsen layer yield reduction across a broad range of plasma conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Dark Energy Survey Year 6 results: Redshift calibration of the MagLim++ lens sample

In this work, we derive and calibrate the redshift distribution of the MagLim++ lens galaxy sample used in the Dark Energy Survey Year 6 (DES Y6) 3 x 2pt cosmology analysis. The 3 x 2pt analysis combines galaxy clustering from the lens galaxy sample and weak gravitational lensing. The redshift distributions are inferred using the SOMPZ method - a Self-Organizing Map framework that combines deep-field multi-band photometry, wide-field data, and a synthetic source injection ( B alrog) catalog. Key improvements over the DES Year 3 (Y3) calibration include a noise-weighted SOM metric, an expanded Balrog catalogue, and an improved scheme for propagating systematic uncertainties, which allows us to generate O(10 8 ) redshift realizations that collectively span the dominant sources of uncertainty. These realizations are then combined with independent clustering-redshift measurements via importance sampling. The resulting calibration achieves typical uncertainties on the mean redshift of 1-2%, corresponding to a 20-30% average reduction relative to DES Y3. We compress the n(z) uncertainties into a small number of orthogonal modes for use in cosmological inference. Marginalizing over these modes leads to only a minor degradation in cosmological constraints. Here, this analysis establishes the MagLim++ sample as a robust lens sample for precision cosmology with DES Y6 and provides a scalable framework for future surveys.

dark energy↗

𝑁-dimensional maximum-entropy tomography via particle sampling

We propose a modified maximum-entropy (MENT) algorithm for six-dimensional phase space tomography. The algorithm uses particle sampling and low-dimensional density estimation to approximate large sets of high-dimensional integrals in the original MENT formulation. We implement this approach using Markov Chain Monte Carlo (MCMC) sampling techniques and demonstrate convergence of six-dimensional MENT on both synthetic and measured data.

Hoover, Austin [Oak Ridge National Laboratory (ORN↗

Code for BALDR Study 07.04

SAND2024-11256O The Code for BALDR Study 07.04 software reproduces results from the BALDR study concerning "Multilabel Proportion Prediction and Out-of-Distribution Detection on Gamma Spectra of Short-Lived Fission Products." This code can reproduce a scientific study following the step numbers present in the file names. The scientific study uses synthetic and measured data to find the best model for the radioisotope proportion estimation task of interest and generates results. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Morrow, Tyler↗