Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

SOC Microstructural Property Estimator

This pre-trained ML model is a tool that uses basic compositional parameters for porous solid oxide cell (SOC) electrodes - the phase fractions and mean particle/pore diameters – as inputs and uses them to estimate additional electrochemical performance parameters: active (i.e., connected) TPB density, all tortuosity factors, and phase pair specific interfacial areas. The electrode is assumed to be composed of two solid phases and a pore phase. The property calculations are performed using neural network regression models trained on a large bank of synthetic electrode microstructural data that NETL has generated using the program DREAM3D (that bank is also hosted on EDX: https://edx.netl.doe.gov/dataset/soc-synthetic-microstructure-bank). This means the generated parameters are based on training from actual measured properties from 3D microstructures, not estimated from geometric simplifications. This tool was developed and is intended to replace percolation theory calculations in models that use hypothetical electrode properties. An example use case would be running SOC performance simulations across a parametric sweep of electrode designs (e.g., varying phase fractions and particle sizes) and assessing how it impacts the electrochemical performance of the SOC. Within the parameter space of the training data (statistics of that parameter space is provided in the readme file), this model achieves sub-5% mean absolute percent errors, an order of magnitude less error than percolation theory across the same parameter space. However, be aware that this tool was developed with parametric simulations in mind, and users are encouraged to assess accuracy for their own specific use case rather than taking accuracy metrics at face value. More info, including a usage guide, is in the included readme file. This tool should be cited with the DOI number provided.

Electrode Microstructure↗

A Physics-Informed Deep Learning Description of Knudsen Layer Reactivity Reduction

A physics-informed neural network (PINN) is used to evaluate the fast ion distribution in the hot spot of an inertial confinement fusion target. The use of tailored input and output layers to the neural network is shown to enable a PINN to learn the parametric solution to the Vlasov–Fokker–Planck equation in the absence of any synthetic or experimental data. As an explicit demonstration of the approach, the specific problem of Knudsen layer fusion yield reduction is treated. Here, the predictions from the Vlasov–Fokker–Planck PINN are used to provide a non-perturbative solution of the fast ion tail in the vicinity of the hot spot, thus allowing the spatial profile of the fusion reactivity to be evaluated for a range of collisionalities and hot spot conditions. Excellent agreement is found between the predictions of the Vlasov–Fokker–Planck PINN and the results from traditional numerical solvers with respect to both the energy and spatial distribution of fast ions and the fusion reactivity profile, demonstrating that the Vlasov–Fokker–Planck PINN provides an accurate and efficient means of determining the impact of Knudsen layer yield reduction across a broad range of plasma conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Dark Energy Survey Year 6 Results: Redshift Calibration of the MagLim++ Lens Sample

In this work, we derive and calibrate the redshift distribution of the MagLim++ lens galaxy sample used in the Dark Energy Survey Year 6 (DES Y6) 3x2pt cosmology analysis. The 3x2pt analysis combines galaxy clustering from the lens galaxy sample and weak gravitational lensing. The redshift distributions are inferred using the SOMPZ method - a Self-Organizing Map framework that combines deep-field multi-band photometry, wide-field data, and a synthetic source injection (Balrog) catalog. Key improvements over the DES Year 3 (Y3) calibration include a noise-weighted SOM metric, an expanded Balrog catalogue, and an improved scheme for propagating systematic uncertainties, which allows us to generate O($10^8$) redshift realizations that collectively span the dominant sources of uncertainty. These realizations are then combined with independent clustering-redshift measurements via importance sampling. The resulting calibration achieves typical uncertainties on the mean redshift of 1-2%, corresponding to a 20-30% average reduction relative to DES Y3. We compress the $n(z)$ uncertainties into a small number of orthogonal modes for use in cosmological inference. Marginalizing over these modes leads to only a minor degradation in cosmological constraints. This analysis establishes the MagLim++ sample as a robust lens sample for precision cosmology with DES Y6 and provides a scalable framework for future surveys.

Giannini, G. [Chicago U., Astron. Astrophys. Ctr.;↗

Generating synthetic signaling networks for in silico modeling studies

Predictive models of signaling pathways have proven to be difficult to develop. Reasons include the uncertainty in the number of species, the complexity in species’ interactions, and the sparseness and uncertainty in experimental data. Traditional approaches to developing mechanistic models rely on collecting experimental data and fitting a single model to that data. This approach works for simple systems but has proven unreliable for complex systems such as biological signaling networks. For example, uncertainty and sparseness of the data often result in overfitted models that have little predictive value beyond recapitulating the experimental data itself. Thus, there is a need to develop new approaches to create predictive mechanistic models of complex systems. However, to determine the effectiveness of any new algorithm, a baseline model is needed to test its performance. To meet this need, we developed a method for generating artificial synthetic networks that are reasonably realistic and thus can be treated as ground truth models. These synthetic models can then be used to generate synthetic data for developing and testing algorithms designed to recover the underlying network topology and associated parameters. Here, we describe a simple approach for generating synthetic signaling networks that can be used for this purpose.

42 ENGINEERING↗

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY↗

Synthetic method of analogues for emerging infectious disease forecasting

The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.

97 MATHEMATICS AND COMPUTING↗

Resource-Adaptive Federated Text Generation with Differential Privacy

In cross-silo federated learning (FL), sensitive text datasets remain confined to local organizations due to privacy regulations, making repeated training for each downstream task both communication-intensive and privacy-demanding. A promising alternative is to generate differentially private (DP) synthetic datasets that approximate the global distribution and can be reused across tasks. However, pretrained large language models (LLMs) often fail under domain shift, and federated finetuning is hindered by computational heterogeneity: only resource-rich clients can update the model, while weaker clients are excluded, amplifying data skew and the adverse effects of DP noise. We propose a flexible participation framework that adapts to client capacities. Strong clients perform DP federated finetuning, while weak clients contribute through a lightweight DP voting mechanism that refines synthetic text. To ensure the synthetic data mirrors the global dataset, we apply control codes (e.g., labels, topics, metadata) that represent each client’s data proportions and constrain voting to semantically coherent subsets. This two-phase approach requires only a single round of communication for weak clients and integrates contributions from all participants. Experiments show that our framework improves distribution alignment and downstream robustness under DP and heterogeneity.

Wang, Jiayi [ORNL]↗

Feedback Controllability Components Analysis (FCCA) v1.0

FCCA is a linear dimensionality reduction method that find subspaces of high-dimensional time-series data that are most feedback controllable. The key innovation is to formulate an objective function that quantifies the joint cost of state reconstruction and state regulation that can be evaluated from purely observational data. To do this, it leverages the duality between controllability and observability. We provide analytic results demonstrating the validity of the cost function. We evaluated this method in both synthetic and real neural data (from multiple organisms and brain areas).

Kumar, Ankit↗

Evaluating the limitations of Bayesian metabolic control analysis

Bayesian Metabolic Control Analysis (BMCA) is a promising framework for inferring metabolic control coefficients in data-limited scenarios, combining Bayesian inference with linear-logarithmic (lin-log) rate laws. These metabolic control coefficients quantify how changes in enzyme activities affect steady-state fluxes and metabolite concentrations across a metabolic network. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCC), and concentration control coefficients (CCC) under varying data availability conditions using three synthetic metabolic network models. We demonstrate that BMCA predictions are highly dependent on the inclusion of flux and enzyme concentration data, with the omission of these datasets leading to severe inaccuracies. In our synthetic, enzyme-perturbation datasets, external metabolite concentrations had minimal impact and, in some cases, their exclusion improved predictions; when external-nutrient perturbations were introduced and those concentrations were observed, gains were at most modest. Additionally, we find that posterior estimation with both ADVI and HMC can underestimate large-magnitude elasticities in our synthetic settings, with ADVI showing somewhat higher variance under strong up-regulation; thus, recovering |elasticity| ≳ 1.5 remains challenging regardless of the inference engine. ADVI also fails to accurately infer allosteric interactions, even when regulatory effects are strong. While BMCA maintains reasonable accuracy in partially recovering the rankings of the highest FCC values, its estimates of absolute values remain constrained by prior assumptions and data limitations. Our findings reveal the BMCA algorithm’s strengths and weaknesses, providing guidance on its application in metabolic engineering, and highlighting the need for methodological refinements to enhance its predictive capabilities.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing Machine Learning and Data Fusion for Accurate Undocumented Well Identification in Satellite Images

This study utilizes satellite data to detect undocumented oil and gas wells, which pose significant environmental concerns, including greenhouse gas emissions. Three key findings emerge from the study. Firstly, the problem of imbalanced data is addressed by recommending oversampling techniques like Rotation–GaussianBlur–Solarization data augmentation (RGS), the Synthetic Minority Over-Sampling Technique (SMOTE), or ADASYN (an extension of SMOTE) over undersampling techniques. The performance of borderline SMOTE is less effective than that of the rest of the oversampling techniques, as its performance relies heavily on the quality and distribution of data near the decision boundary. Secondly, incorporating pre-trained models trained on large-scale datasets enhances the models’ generalization ability, with models trained on one county’s dataset demonstrating high overall accuracy, recall, and F1 scores that can be extended to other areas. This transferability of models allows for wider application. Lastly, including persistent homology (PH) as an additional input improves performance for in-distribution testing but may affect the model’s generalization for out-of-distribution testing. A careful consideration of PH’s impact on overall performance and generalizability is recommended. Overall, this study provides a robust approach to identifying undocumented oil and gas wells, contributing to the acceleration of a net-zero economy and supporting environmental sustainability efforts.

SMOTE↗

Evaluating the limitations of Bayesian metabolic control analysis

AbstractBayesian Metabolic Control Analysis (BMCA) has emerged as a promising framework for inferring metabolic control coefficients in data-limited scenarios by integrating Bayesian inference with linlog rate laws. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCCs), and concentration control coefficients (CCCs) under varying data availability conditions using three synthetic metabolic network models. Our findings highlight the strengths and weaknesses of BMCA, guiding its application in metabolic engineering and emphasizing the need for methodological refinements.Author summaryUnderstanding how enzymes control metabolic pathways is crucial for optimizing biomanufacturing and synthetic biology applications. Bayesian Metabolic Control Analysis (BMCA) is a promising computational method that integrates Bayesian inference with metabolic control analysis to estimate key control parameters, even in cases with limited experimental data. However, the accuracy and limitations of BMCA remain unclear. In this study, we systematically evaluate BMCA using three synthetic metabolic networks to determine how different types of physiological data impact its predictive performance. We find that BMCA requires flux and enzyme concentration data for accurate predictions, while external metabolite concentrations contribute little. Additionally, BMCA fails to predict elasticity values beyond a magnitude of 1.5 and reliably infer allosteric regulation, even when strong regulatory interactions exist. In addition, BMCA does not accurately rank metabolic control points, which may limit its utility in identifying key enzymes in engineered pathways. Our work provides practical insights into when and how BMCA can be applied, guiding future research in metabolic modeling and control analysis.

Shin, Janis (ORCID:0000000216572455)↗

Multimodal super-resolution: discovering hidden physics and its application to fusion plasmas

Understanding complex physical systems often requires integrating data from multiple diagnostics, each with limited resolution or coverage. We present a machine learning framework that reconstructs synthetic high-temporal-resolution data for a target diagnostic using information from other diagnostics, without direct target measurements during the inference. This multimodal super-resolution technique improves diagnostic robustness and enables monitoring even in case of measurement failures or degradation. Applied to fusion plasmas, our method targets edge-localized modes (ELMs), which can damage plasma-facing materials. By reconstructing super-resolution Thomson Scattering data from complementary diagnostics, we uncover fine-scale plasma dynamics and validate the role of resonant magnetic perturbations (RMPs) in ELM suppression through magnetic island formation. The approach provides new observation supporting the plasma profile flattening due to these islands. Our results demonstrate the framework’s ability to generate high-fidelity synthetic diagnostics, offering a powerful tool for ELM control development in future reactors like ITER. The approach is broadly transferable to other domains facing sparse, incomplete, or degraded diagnostic data, opening new avenues for discovery.

Jalalvand, Azarakhsh [Princeton Univ., NJ (United ↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Chemistry↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Arasteh, Kourosh [Lawrence Livermore National Labo↗

Updimensioning strategy derived from synthetic equiaxed grain structures for approximating 3D grain size distributions from 2D visualizations with 1D parameters

We generated synthetic equiaxed grain structures using computer graphics software to explore the relationship between various grain size determination methods and true three-dimensional (3D) grain diameters. Mirroring grain measurement techniques, the synthetic 3D grain structures are imaged as 2D micrographs which are measured to yield 1D grain size parameters. Synthetic grain structures provide data at a mass scale and permit exploration of both polished and fractured surface micrographs, revealing one-to-one correspondence between exposed 2D grain cross-sections and individual 3D grains. Analysis of this correspondence yielded a procedure to approximate 3D equiaxed grain size and volume distributions based on the mode of the 2D fractograph grain size distribution. The 3D approximation procedure is shown to be less susceptible to different imaging conditions that affect small, undiscernible grains compared to the standard planimetric and linear intercept methods, which by design also tend to underestimate the 3D grain diameter. The procedure requires larger sample sizes to lower variance and a deeper analysis which could become more practical with machine learning (ML) models for grain boundary segmentation, which synthetic grain structures can help train. This work lays the foundation for analyzing other grain distributions such as columnar and composite grains in similar depth.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Redox‐Active Frustrated Lewis Pair‐Mediated B—H Bond Activation: From Proton Transfer to THF Ring Opening

The CAAC-stabilized dithiolene (L 0 ) zwitterion (1), an unusual redox-active intramolecular frustrated Lewis pair (FLP), activates the B─H bond of boranes via hydride-coupled reverse electron transfer processes. The reactions of 1 with catecholborane in THF give a zwitterionic bis(dithiolene)-based spiroborate ( 2 ), in which one sulphur atom (at the C2 carbon) bonds to the (CH 2 ) 4 OB(O) 2 C 6 H 4 chain due to the catecholborane (CatBH)/S thiourea Lewis pair-mediated THF ring opening. In addition to 2 , [CAAC(H)] + [(Cat) 2 B] − , ( 3 ), CAAC(H) 2 , ( 4 ), and [CAAC(H) 3 ] + [(Cat) 2 B] − , ( 5 ), are also isolated from these reactions. In addition to synthetic, structural, and spectroscopic data, a plausible reaction mechanism is proposed. This finding provides compelling experimental evidence of FLP-mediated B─H activation via net proton transfer.

carbenes↗

High-resolution fully-polarimetric synthetic aperture radar dataset

Fully-polarimetric synthetic aperture radar (PolSAR) data contain a rich body of elementary scattering physics information that is critically valuable for a broad range of applications and scientific purposes. However, there is a lack of available high-resolution (< 0.3048-m) data available for PolSAR phenomenology research. This article introduces a high-resolution PolSAR data set collected and provided by Sandia National Laboratories (SNL). The data sets were collected to support studying high-resolution scattering physics from different types of clutter and applications such as polarimetric-based terrain classification.

West, Roger Derek↗

Drought-induced changes in groundwater-surface water exchange at Lake Mead area

This study focuses on the Lake Mead region in the southwestern United States, a key water reservoir serving over 25 million people and agricultural lands across several states. The area has experienced recurring anthropogenic droughts since the early 2000s. We investigate the hydrological response of the coupled surface-groundwater system to the 2020–2022 drought, one of the most severe on record. To this end, we use Sentinel-1 Interferometric Synthetic Aperture Radar (InSAR) data to quantify vertical land motion caused by the elastic response of the crust to water-mass loss in the lake vicinity. Next, we apply an inverse elastic load modeling framework to quantify water loss. We further assess possible hydraulic connectivity between Lake Mead and adjacent groundwater reservoirs. We detect ground uplift of up to 8 mm/yr near the lake center, likely due to crustal rebound from reduced water-mass loading. We estimated the total water storage loss at 3.03 ± 0.25 km 3 /yr across a 3150 km 2 area surrounding Lake Mead. Groundwater accounts for approximately a third of that, being 0.94 ± 0.32 km 3 /yr. In addition, observed time lags of 6–98 days between lake and groundwater level responses, corresponding to a lateral diffusivity of 3.2–86 m 2 /s, suggest spatially variable connectivity between the lake and aquifers. These findings highlight that drought impacts propagate through the subsurface within interconnected systems, resulting in reduced buffering capacity of groundwater resources following droughts, and emphasizing the need for more integrated surface and groundwater management strategies to enhance resilience under climate and anthropogenic stressors.

58 GEOSCIENCES↗