Search NASA⌕ Search

SEARCH · Search NASA

Results for “Synthetic data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Generating synthetic signaling networks for in silico modeling studies

Predictive models of signaling pathways have proven to be difficult to develop. Reasons include the uncertainty in the number of species, the complexity in species’ interactions, and the sparseness and uncertainty in experimental data. Traditional approaches to developing mechanistic models rely on collecting experimental data and fitting a single model to that data. This approach works for simple systems but has proven unreliable for complex systems such as biological signaling networks. For example, uncertainty and sparseness of the data often result in overfitted models that have little predictive value beyond recapitulating the experimental data itself. Thus, there is a need to develop new approaches to create predictive mechanistic models of complex systems. However, to determine the effectiveness of any new algorithm, a baseline model is needed to test its performance. To meet this need, we developed a method for generating artificial synthetic networks that are reasonably realistic and thus can be treated as ground truth models. These synthetic models can then be used to generate synthetic data for developing and testing algorithms designed to recover the underlying network topology and associated parameters. Here, we describe a simple approach for generating synthetic signaling networks that can be used for this purpose.

42 ENGINEERING↗

Machine Learning‐Assisted Microearthquake Location Workflow for Monitoring the Newberry Enhanced Geothermal System

Abstract Enhanced geothermal systems (EGS) offer a sustainable energy source but face challenges in accurately locating microearthquakes induced during reservoir stimulation. Locating these microearthquakes provides reliable feedback on the stimulation progress. Current deep learning methods for locating earthquakes require extensive data sets for training, which is problematic as detected microearthquakes are often limited. To address the scarcity of training data, we propose a practical workflow using probabilistic multilayer perceptron (PMLP) which predicts microearthquake locations from cross‐correlation time lags in waveforms. Utilizing a 3D velocity model of Newberry site derived from ambient noise interferometry, we generate numerous synthetic microearthquakes and 3D acoustic waveforms for PMLP training. Accurate synthetic tests prompt us to apply the trained network to the 2012 and 2014 stimulation field waveforms. To enhance the accuracy of source localization, we carefully handpick the P‐arrival times. Predictions on the 2012 stimulation data set show major microseismic activity at depths of 0.5–1.2 km, correlating with a known casing leakage scenario. In the 2014 data set, the majority of predictions concentrate at 2.0–2.9 km depths, consistent with results obtained from conventional physics‐based inversion, and align with the presence of natural fractures from 2.0 to 2.7 km. We validate our findings by comparing the synthetic and field picks, demonstrating a satisfactory match for the first arrivals. By combining the benefits of quick inference speeds and accurate location predictions, we demonstrate the feasibility of using realistic synthetic data set to locate microseismicity for EGS monitoring.

15 GEOTHERMAL ENERGY↗

Synthetic method of analogues for emerging infectious disease forecasting

The Method of Analogues (MOA) has gained popularity in the past decade for infectious disease forecasting due to its non-parametric nature. In MOA, the local behavior observed in a time series is matched to the local behaviors of several historical time series. The known values that directly follow the historical time series that best match the observed time series are used to calculate a forecast. This non-parametric approach leverages historical trends to produce forecasts without extensive parameterization, making it highly adaptable. However, MOA is limited in scenarios where historical data is sparse. This limitation was particularly evident during the early stages of the COVID-19 pandemic, where the emerging global epidemic had little-to-no historical data. In this work, we propose a new method inspired by MOA, called the Synthetic Method of Analogues (sMOA). sMOA replaces historical disease data with a library of synthetic data that describe a broad range of possible disease trends. This model circumvents the need to estimate explicit parameter values by instead matching segments of ongoing time series data to a comprehensive library of synthetically generated segments of time series data. We demonstrate that sMOA has competitive performance with state-of-the-art infectious disease forecasting models, out-performing 78% of models from the COVID-19 Forecasting Hub in terms of averaged Mean Absolute Error and 76% of models from the COVID-19 Forecasting Hub in terms of averaged Weighted Interval Score. Additionally, we introduce a novel uncertainty quantification methodology designed for the onset of emerging epidemics. Developing versatile approaches that do not rely on historical data and can maintain high accuracy in the face of novel pandemics is critical for enhancing public health decision-making and strengthening preparedness for future outbreaks.

97 MATHEMATICS AND COMPUTING↗

teemi: An open-source literate programming approach for iterative design-build-test-learn cycles in bioengineering

Synthetic biology dictates the data-driven engineering of biocatalysis, cellular functions, and organism behavior. Integral to synthetic biology is the aspiration to efficiently find, access, interoperate, and reuse high-quality data on genotype-phenotype relationships of native and engineered biosystems under FAIR principles, and from this facilitate forward-engineering strategies. However, biology is complex at the regulatory level, and noisy at the operational level, thus necessitating systematic and diligent data handling at all levels of the design, build, and test phases in order to maximize learning in the iterative design-build-test-learn engineering cycle. To enable user-friendly simulation, organization, and guidance for the engineering of biosystems, we have developed an open-source python-based computer-aided design and analysis platform operating under a literate programming user-interface hosted on Github. The platform is called teemi and is fully compliant with FAIR principles. In this study we apply teemi for i) designing and simulating bioengineering, ii) integrating and analyzing multivariate datasets, and iii) machine-learning for predictive engineering of metabolic pathway designs for production of a key precursor to medicinal alkaloids in yeast. The teemi platform is publicly available at PyPi and GitHub.

59 BASIC BIOLOGICAL SCIENCES↗

Resource-Adaptive Federated Text Generation with Differential Privacy

In cross-silo federated learning (FL), sensitive text datasets remain confined to local organizations due to privacy regulations, making repeated training for each downstream task both communication-intensive and privacy-demanding. A promising alternative is to generate differentially private (DP) synthetic datasets that approximate the global distribution and can be reused across tasks. However, pretrained large language models (LLMs) often fail under domain shift, and federated finetuning is hindered by computational heterogeneity: only resource-rich clients can update the model, while weaker clients are excluded, amplifying data skew and the adverse effects of DP noise. We propose a flexible participation framework that adapts to client capacities. Strong clients perform DP federated finetuning, while weak clients contribute through a lightweight DP voting mechanism that refines synthetic text. To ensure the synthetic data mirrors the global dataset, we apply control codes (e.g., labels, topics, metadata) that represent each client’s data proportions and constrain voting to semantically coherent subsets. This two-phase approach requires only a single round of communication for weak clients and integrates contributions from all participants. Experiments show that our framework improves distribution alignment and downstream robustness under DP and heterogeneity.

Wang, Jiayi [ORNL]↗

Feedback Controllability Components Analysis (FCCA) v1.0

FCCA is a linear dimensionality reduction method that find subspaces of high-dimensional time-series data that are most feedback controllable. The key innovation is to formulate an objective function that quantifies the joint cost of state reconstruction and state regulation that can be evaluated from purely observational data. To do this, it leverages the duality between controllability and observability. We provide analytic results demonstrating the validity of the cost function. We evaluated this method in both synthetic and real neural data (from multiple organisms and brain areas).

Kumar, Ankit↗

Evaluating the limitations of Bayesian metabolic control analysis

Bayesian Metabolic Control Analysis (BMCA) is a promising framework for inferring metabolic control coefficients in data-limited scenarios, combining Bayesian inference with linear-logarithmic (lin-log) rate laws. These metabolic control coefficients quantify how changes in enzyme activities affect steady-state fluxes and metabolite concentrations across a metabolic network. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCC), and concentration control coefficients (CCC) under varying data availability conditions using three synthetic metabolic network models. We demonstrate that BMCA predictions are highly dependent on the inclusion of flux and enzyme concentration data, with the omission of these datasets leading to severe inaccuracies. In our synthetic, enzyme-perturbation datasets, external metabolite concentrations had minimal impact and, in some cases, their exclusion improved predictions; when external-nutrient perturbations were introduced and those concentrations were observed, gains were at most modest. Additionally, we find that posterior estimation with both ADVI and HMC can underestimate large-magnitude elasticities in our synthetic settings, with ADVI showing somewhat higher variance under strong up-regulation; thus, recovering |elasticity| ≳ 1.5 remains challenging regardless of the inference engine. ADVI also fails to accurately infer allosteric interactions, even when regulatory effects are strong. While BMCA maintains reasonable accuracy in partially recovering the rankings of the highest FCC values, its estimates of absolute values remain constrained by prior assumptions and data limitations. Our findings reveal the BMCA algorithm’s strengths and weaknesses, providing guidance on its application in metabolic engineering, and highlighting the need for methodological refinements to enhance its predictive capabilities.

59 BASIC BIOLOGICAL SCIENCES↗

Harnessing Machine Learning and Data Fusion for Accurate Undocumented Well Identification in Satellite Images

This study utilizes satellite data to detect undocumented oil and gas wells, which pose significant environmental concerns, including greenhouse gas emissions. Three key findings emerge from the study. Firstly, the problem of imbalanced data is addressed by recommending oversampling techniques like Rotation–GaussianBlur–Solarization data augmentation (RGS), the Synthetic Minority Over-Sampling Technique (SMOTE), or ADASYN (an extension of SMOTE) over undersampling techniques. The performance of borderline SMOTE is less effective than that of the rest of the oversampling techniques, as its performance relies heavily on the quality and distribution of data near the decision boundary. Secondly, incorporating pre-trained models trained on large-scale datasets enhances the models’ generalization ability, with models trained on one county’s dataset demonstrating high overall accuracy, recall, and F1 scores that can be extended to other areas. This transferability of models allows for wider application. Lastly, including persistent homology (PH) as an additional input improves performance for in-distribution testing but may affect the model’s generalization for out-of-distribution testing. A careful consideration of PH’s impact on overall performance and generalizability is recommended. Overall, this study provides a robust approach to identifying undocumented oil and gas wells, contributing to the acceleration of a net-zero economy and supporting environmental sustainability efforts.

SMOTE↗

Evaluating the limitations of Bayesian metabolic control analysis

AbstractBayesian Metabolic Control Analysis (BMCA) has emerged as a promising framework for inferring metabolic control coefficients in data-limited scenarios by integrating Bayesian inference with linlog rate laws. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCCs), and concentration control coefficients (CCCs) under varying data availability conditions using three synthetic metabolic network models. Our findings highlight the strengths and weaknesses of BMCA, guiding its application in metabolic engineering and emphasizing the need for methodological refinements.Author summaryUnderstanding how enzymes control metabolic pathways is crucial for optimizing biomanufacturing and synthetic biology applications. Bayesian Metabolic Control Analysis (BMCA) is a promising computational method that integrates Bayesian inference with metabolic control analysis to estimate key control parameters, even in cases with limited experimental data. However, the accuracy and limitations of BMCA remain unclear. In this study, we systematically evaluate BMCA using three synthetic metabolic networks to determine how different types of physiological data impact its predictive performance. We find that BMCA requires flux and enzyme concentration data for accurate predictions, while external metabolite concentrations contribute little. Additionally, BMCA fails to predict elasticity values beyond a magnitude of 1.5 and reliably infer allosteric regulation, even when strong regulatory interactions exist. In addition, BMCA does not accurately rank metabolic control points, which may limit its utility in identifying key enzymes in engineered pathways. Our work provides practical insights into when and how BMCA can be applied, guiding future research in metabolic modeling and control analysis.

Shin, Janis (ORCID:0000000216572455)↗

Multimodal super-resolution: discovering hidden physics and its application to fusion plasmas

Understanding complex physical systems often requires integrating data from multiple diagnostics, each with limited resolution or coverage. We present a machine learning framework that reconstructs synthetic high-temporal-resolution data for a target diagnostic using information from other diagnostics, without direct target measurements during the inference. This multimodal super-resolution technique improves diagnostic robustness and enables monitoring even in case of measurement failures or degradation. Applied to fusion plasmas, our method targets edge-localized modes (ELMs), which can damage plasma-facing materials. By reconstructing super-resolution Thomson Scattering data from complementary diagnostics, we uncover fine-scale plasma dynamics and validate the role of resonant magnetic perturbations (RMPs) in ELM suppression through magnetic island formation. The approach provides new observation supporting the plasma profile flattening due to these islands. Our results demonstrate the framework’s ability to generate high-fidelity synthetic diagnostics, offering a powerful tool for ELM control development in future reactors like ITER. The approach is broadly transferable to other domains facing sparse, incomplete, or degraded diagnostic data, opening new avenues for discovery.

Jalalvand, Azarakhsh [Princeton Univ., NJ (United ↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Chemistry↗

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Arasteh, Kourosh [Lawrence Livermore National Labo↗

Updimensioning strategy derived from synthetic equiaxed grain structures for approximating 3D grain size distributions from 2D visualizations with 1D parameters

We generated synthetic equiaxed grain structures using computer graphics software to explore the relationship between various grain size determination methods and true three-dimensional (3D) grain diameters. Mirroring grain measurement techniques, the synthetic 3D grain structures are imaged as 2D micrographs which are measured to yield 1D grain size parameters. Synthetic grain structures provide data at a mass scale and permit exploration of both polished and fractured surface micrographs, revealing one-to-one correspondence between exposed 2D grain cross-sections and individual 3D grains. Analysis of this correspondence yielded a procedure to approximate 3D equiaxed grain size and volume distributions based on the mode of the 2D fractograph grain size distribution. The 3D approximation procedure is shown to be less susceptible to different imaging conditions that affect small, undiscernible grains compared to the standard planimetric and linear intercept methods, which by design also tend to underestimate the 3D grain diameter. The procedure requires larger sample sizes to lower variance and a deeper analysis which could become more practical with machine learning (ML) models for grain boundary segmentation, which synthetic grain structures can help train. This work lays the foundation for analyzing other grain distributions such as columnar and composite grains in similar depth.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Redox‐Active Frustrated Lewis Pair‐Mediated B—H Bond Activation: From Proton Transfer to THF Ring Opening

The CAAC-stabilized dithiolene (L 0 ) zwitterion (1), an unusual redox-active intramolecular frustrated Lewis pair (FLP), activates the B─H bond of boranes via hydride-coupled reverse electron transfer processes. The reactions of 1 with catecholborane in THF give a zwitterionic bis(dithiolene)-based spiroborate ( 2 ), in which one sulphur atom (at the C2 carbon) bonds to the (CH 2 ) 4 OB(O) 2 C 6 H 4 chain due to the catecholborane (CatBH)/S thiourea Lewis pair-mediated THF ring opening. In addition to 2 , [CAAC(H)] + [(Cat) 2 B] − , ( 3 ), CAAC(H) 2 , ( 4 ), and [CAAC(H) 3 ] + [(Cat) 2 B] − , ( 5 ), are also isolated from these reactions. In addition to synthetic, structural, and spectroscopic data, a plausible reaction mechanism is proposed. This finding provides compelling experimental evidence of FLP-mediated B─H activation via net proton transfer.

carbenes↗

High-resolution fully-polarimetric synthetic aperture radar dataset

Fully-polarimetric synthetic aperture radar (PolSAR) data contain a rich body of elementary scattering physics information that is critically valuable for a broad range of applications and scientific purposes. However, there is a lack of available high-resolution (< 0.3048-m) data available for PolSAR phenomenology research. This article introduces a high-resolution PolSAR data set collected and provided by Sandia National Laboratories (SNL). The data sets were collected to support studying high-resolution scattering physics from different types of clutter and applications such as polarimetric-based terrain classification.

West, Roger Derek↗

Drought-induced changes in groundwater-surface water exchange at Lake Mead area

This study focuses on the Lake Mead region in the southwestern United States, a key water reservoir serving over 25 million people and agricultural lands across several states. The area has experienced recurring anthropogenic droughts since the early 2000s. We investigate the hydrological response of the coupled surface-groundwater system to the 2020–2022 drought, one of the most severe on record. To this end, we use Sentinel-1 Interferometric Synthetic Aperture Radar (InSAR) data to quantify vertical land motion caused by the elastic response of the crust to water-mass loss in the lake vicinity. Next, we apply an inverse elastic load modeling framework to quantify water loss. We further assess possible hydraulic connectivity between Lake Mead and adjacent groundwater reservoirs. We detect ground uplift of up to 8 mm/yr near the lake center, likely due to crustal rebound from reduced water-mass loading. We estimated the total water storage loss at 3.03 ± 0.25 km 3 /yr across a 3150 km 2 area surrounding Lake Mead. Groundwater accounts for approximately a third of that, being 0.94 ± 0.32 km 3 /yr. In addition, observed time lags of 6–98 days between lake and groundwater level responses, corresponding to a lateral diffusivity of 3.2–86 m 2 /s, suggest spatially variable connectivity between the lake and aquifers. These findings highlight that drought impacts propagate through the subsurface within interconnected systems, resulting in reduced buffering capacity of groundwater resources following droughts, and emphasizing the need for more integrated surface and groundwater management strategies to enhance resilience under climate and anthropogenic stressors.

58 GEOSCIENCES↗

Forest aboveground biomass estimation through integration of sentinel-2 and PALSAR-2 time series: assessing models trained on GEDI and field inventory benchmarks

Accurate and spatially explicit forest Aboveground Biomass (AGB) mapping through remote sensing is critical for quantifying terrestrial carbon stocks and informing effective forest management strategies. However, AGB estimation in dense forests with complex terrain remains challenging due to satellite sensor signal saturation problem (saturation issue occurs in high biomass forests), structural complexity, and limited ground truth for calibration. This study presents a novel framework that integrates multi-temporal Sentinel-2 optical imagery, ALOS PALSAR-2 Synthetic Aperture Radar (SAR) data, and topographic variables with explainable Machine Learning to map AGB across mountainous forests within subtropical and temperate oceanic climate zones of Mexico. We evaluate the effects of temporal granularity and sensor synergy by comparing multiple temporal inputs and sensor configurations (Sentinel-2, PALSAR-2, and their fusion), and assess model performance using two reference datasets: NASA GEDI LiDAR-derived biomass and Mexico’s National Forest and Soil Inventory (INFyS). Our results showed that models trained on INFyS consistently outperformed those trained on GEDI, highlighting limitations in GEDI’s reliability in biomass estimates within this study region. Furthermore, the integration of Sentinel-2 and PALSAR-2 provided improved predictions compared to single-sensor models, particularly when combined with temporally explicit yearly statistics. The best-performing model, which was trained on INFyS data, and considered both Sentinel-2 and PALSAR-2 yearly statistics, as well as topographic variables, achieved an R2 of 0.64, RMSE of 51.10 Mg/ha, and relative RMSE (rRMSE) of 58.69%. Explainable ML analysis identified Sentinel-2 spectral indices and topographic features as key predictors, while PALSAR-2 metrics provided complementary information, partially mitigating saturation effects in high-biomass areas. Specifically, integrating both sensors substantially improved AGB estimation in high biomass forest (≥200 Mg/ha), yielding 98% gains over optical-only model, with resulting estimates exceeding GEDI L4B by 29% and ESA-CCI-BIOMASS by 174%. Terrain-stratified analysis indicated close agreement with GEDI in low-slope areas, with increasing divergence as slope steepness increased, while estimates remained consistently higher than ESA-CCI-BIOMASS across all slope classes. The proposed approach advances multi-sensor fusion and temporal feature engineering for AGB mapping using open-access satellite datasets, providing a scalable and reproducible framework for annual biomass monitoring in topographically complex mountainous forests. The resulting 25 m resolution biomass product has the potential to provide spatially detailed information for forest monitoring and may support applications in carbon accounting and forest management.

54 ENVIRONMENTAL SCIENCES↗

Next generation Arctic vegetation maps: Aboveground plant biomass and woody dominance mapped at 30 m resolution across the tundra biome

The Arctic is warming faster than anywhere else on Earth, placing tundra ecosystems at the forefront of global climate change. Plant biomass is a fundamental ecosystem attribute that is sensitive to changes in climate, closely tied to ecological function, and crucial for constraining ecosystem carbon dynamics. However, the amount, functional composition, and distribution of plant biomass are only coarsely quantified across the Arctic. Therefore, we developed the first moderate resolution (30 m) maps of live aboveground plant biomass (g m −2 ) and woody plant dominance (%) for the Arctic tundra biome, including the mountainous Oro Arctic. We modeled biomass for the year 2020 using a new synthesis dataset of field biomass harvest measurements, Landsat satellite seasonal synthetic composites, ancillary geospatial data, and machine learning models. Additionally, we quantified pixel-wise uncertainty in biomass predictions using Monte Carlo simulations and validated the models using a robust, spatially blocked and nested cross-validation procedure. Observed plant and woody plant biomass values ranged from 0 to ∼6000 g m −2 (mean ≈ 350 g m −2 ), while predicted values ranged from 0 to ∼4000 g m −2 (mean ≈ 275 g m −2 ), resulting in model validation root-mean-squared-error (RMSE) ≈ 400 g m −2 and R 2 ≈ 0.6. Our maps not only capture large-scale patterns of plant biomass and woody plant dominance across the Arctic that are linked to climatic variation (e.g., thawing degree days), but also illustrate how fine-scale patterns are shaped by local surface hydrology, topography, and past disturbance. By providing data on plant biomass across Arctic tundra ecosystems at the highest resolution to date, our maps can significantly advance research and inform decision-making on topics ranging from Arctic vegetation monitoring and wildlife conservation to carbon accounting and land surface modeling.

Climate change↗