Search NASA⌕ Search

SEARCH · Search NASA

Results for “data noise”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Dark Energy Survey Year 3 results: Simulation-based cosmological inference with wavelet harmonics, scattering transforms, and moments of weak lensing mass maps. II. cosmological results

Here, we present a simulation-based cosmological analysis using a combination of Gaussian and non-Gaussian statistics of the weak lensing mass (convergence) maps from the first three years of the Dark Energy Survey. We implement the following: (1) second and third moments; (2) wavelet phase harmonics; (3) the scattering transform. Our analysis is fully based on simulations, spans a space of seven 𝑤 Cold Dark Matter (𝑤⁢ CDM) cosmological parameters, and forward models the most relevant sources of systematics inherent in the data: masks, noise variations, clustering of the sources, intrinsic alignments, and shear and redshift calibration. We implement a neural network compression of the summary statistics, and we estimate the parameter posteriors using a simulation-based inference approach. Including and combining different non-Gaussian statistics is a powerful tool that strongly improves constraints over Gaussian statistics (in our case, the second moments); in particular, the figure of merit (𝑆 8 , Ω m ) is improved by 70% (Λ ⁢CDM) and 90% (𝑤 ⁢CDM). When all the summary statistics are combined, we achieve a 2% constraint on the amplitude of fluctuations parameter 𝑆 8 ≡ 𝜎 8 ⁢(Ω m /0.3) 0.5 , obtaining 𝑆 8 = 0.794 ±0.017 (Λ⁢ CDM) and 𝑆 8 = 0.817 ±0.021 (𝑤 ⁢CDM), and a ∼10% constraint on Ω m , obtaining Ω m =0.259 ±0.025 (Λ ⁢CDM) and Ω m = 0.273 ±0.029 (𝑤⁢ CDM). In the context of the 𝑤⁢ CDM scenario, these statistics also strengthen the constraints on the parameter 𝑤, obtaining 𝑤 <−0.72. The constraints from different statistics are shown to be internally consistent (with a 𝑝-value>0.1 for all combinations of statistics examined). We compare our results to other weak lensing results from the first three years of the Dark Energy Survey data, finding good consistency; we also compare with results from external datasets, such as planck constraints from the cosmic microwave background, finding statistical agreement, with discrepancies no greater than <2.2⁢𝜎.

79 ASTRONOMY AND ASTROPHYSICS↗

Beam Loss Assessment Through Use of Photomultiplier Tubes

Modern accelerators aim to deliver maximal beam current at stable energy with minimal beam loss. Environmental changes, among other factors, can result in increased beam loss and decreased beam throughput, prompting daily retuning of the accelerator. The compact nature of the oldest part of the linear accelerator limits the available beam instrumentation, making beam loss assessment and tuning difficult. Thus, additional devices for beam loss monitoring must be considered. Photomultiplier tube-based beam loss monitors (BLMs) were installed along the Fermilab drift tube Linac to assess beam loss. Due to noise, the data from the installed photomultiplier tubes was difficult to assess. After noise reduction and signal analysis, it was found that the signals produced by the photomultiplier tubes in response to beam loss were consistent for a given configuration and therefore a reasonable measure of beam loss. This project lays groundwork for future work in beam loss assessment using photomultiplier tubes, with the automation of the process developed in this project being the next step in this effort.

Waggoner, Alexander↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Denoising of imaginary time response functions with Hankel projections

Imaginary-time response functions of finite-temperature quantum systems are often obtained with methods that exhibit stochastic or systematic errors. Reducing these errors comes at a large computational cost—in quantum Monte Carlo simulations, the reduction of noise by a factor of two incurs a simulation cost of a factor of four. In this paper, we relate certain imaginary-time response functions to an inner product on the space of linear operators on Fock space. We then show that data with noise typically does not respect the positive definiteness of its associated Gramian. The Gramian has the structure of a Hankel matrix. As a method for denoising noisy data, we introduce an alternating projection algorithm that finds the closest positive definite Hankel matrix consistent with noisy data. We test our methodology at the example of fermion Green's functions for continuous-time quantum Monte Carlo data and show remarkable improvements of the error, reducing noise by a factor of up to 20 in practical examples. We argue that Hankel projections should be used whenever finite-temperature imaginary-time data of response functions with errors is analyzed, be it in the context of quantum Monte Carlo, quantum computing, or in approximate semianalytic methodologies. Published by the American Physical Society 2024

Yu, Yang (ORCID:0000000186178878)↗

Optimising the processing and storage of visibilities using lossy compression

The next-generation radio astronomy instruments are providing a massive increase in sensitivity and coverage, largely through increasing the number of stations in the array and the frequency span sampled. The two primary problems encountered when processing the resultant avalanche of data are the need for abundant storage and the constraints imposed by I/O, as I/O bandwidths drop significantly on cold storage. An example of this is the data deluge expected from the SKA Telescopes of more than 60 PB per day, all to be stored on the buffer filesystem. While compressing the data is an obvious solution, the impacts on the final data products are hard to predict. In this paper, we chose an error-controlled compressor – MGARD – and applied it to simulated SKA-Mid and real pathfinder visibility data, in noise-free and noise-dominated regimes. As the data have an implicit error level in the system temperature, using an error bound in compression provides a natural metric for compression. MGARD ensures the compression incurred errors adhere to the user-prescribed tolerance. To measure the degradation of images reconstructed using the lossy compressed data, we proposed a list of diagnostic measures, exploring the trade-off between these error bounds and the corresponding compression ratios, as well as the impact on science quality derived from the lossy compressed data products through a series of experiments. We studied the global and local impacts on the output images for continuum and spectral line examples. We found relative error bounds of as much as 10%, which provide compression ratios of about 20, have a limited impact on the continuum imaging as the increased noise is less than the image RMS, whereas a 1% error bound (compression ratio of 8) introduces an increase in noise of about an order of magnitude less than the image RMS. For extremely sensitive observations and for very precious data, we would recommend a 0.1% error bound with compression ratios of about 4. These have noise impacts two orders of magnitude less than the image RMS levels. At these levels, the limits are due to instabilities in the deconvolution methods. We compared the results to the alternative compression tool DYSCO, in both the impacts on the images and in the relative flexibility. MGARD provides better compression for similar error bounds and has a host of potentially powerful additional features.

Techniques: interferometric↗

Neural Network Prediction of Strong Lensing Systems with Domain Adaptation and Uncertainty Quantification

Modeling strong gravitational lenses is computationally expensive for the complex data from modern and next-generation cosmic surveys. Deep learning has emerged as a promising approach for finding lenses and predicting lensing parameters, such as the Einstein radius. Mean-variance Estimators (MVEs) are a common approach for obtaining aleatoric (data) uncertainties from a neural network prediction. However, neural networks have not been demonstrated to perform well on out-of-domain target data successfully - e.g., when trained on simulated data and applied to real, observational data. In this work, we perform the first study of the efficacy of MVEs in combination with unsupervised domain adaptation (UDA) on strong lensing data. The source domain data is noiseless, and the target domain data has noise mimicking modern cosmology surveys. We find that adding UDA to MVE increases the accuracy on the target data by a factor of about two over an MVE model without UDA. Including UDA also permits much more well-calibrated aleatoric uncertainty predictions. Advancements in this approach may enable future applications of MVE models to real observational data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Finding Real Uncertainties From Physical Simulations

Modeling strong gravitational lenses is computationally expensive for the complex data from modern and next-generation cosmic surveys. Deep learning has emerged as a promising approach for finding lenses and predicting lensing parameters, such as the Einstein radius. Mean-variance Estimators (MVEs) are a common approach for obtaining aleatoric (data) uncertainties from a neural network prediction. However, neural networks have not been demonstrated to perform well on out-of-domain target data successfully - e.g., when trained on simulated data and applied to real, observational data. In this work, we perform the first study of the efficacy of MVEs in combination with unsupervised domain adaptation (UDA) on strong lensing data. The source domain data is noiseless, and the target domain data has noise mimicking modern cosmology surveys. We find that adding UDA to MVE increases the accuracy on the target data by a factor of about two over an MVE model without UDA. Including UDA also permits much more well-calibrated aleatoric uncertainty predictions. Advancements in this approach may enable future applications of MVE models to real observational data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Basic Research Needs for Inverse Methods for Complex Systems under Uncertainty

Inverse problems, which aim to infer unknown properties of a system using experimental and observational data, are central to addressing many of the U.S. Department of Energy’s (DOE) most critical scientific and engineering challenges. Accurate, computationally efficient, and data-efficient solutions to inverse problems are essential for advancing DOE mission-critical science drivers, including analyzing data from large-scale experimental facilities, optimizing fusion reactor performance, accelerating materials discovery, enhancing geophysical imaging, improving wildfire predictions, and enabling autonomous systems and digital twins. However, these problems are becoming increasingly complex, often involving nonlinear, highdimensional, and interconnected systems and models that span multiple physics and scales, while relying on data with varying quantity, quality, and information content. Compounding these challenges is the uncertainty inherent in DOE-relevant systems, where errors in inputs, noise in data, incompleteness of data, and discrepancies between models and reality constrain the accuracy and precision of solutions. At the same time, the convergence of recent scientific computing trends—scientific machine learning, artificial intelligence, and computing advances such as exascale computing—is creating unprecedented opportunities for tackling these challenges. The cross-cutting nature of inverse problems, combined with their growing complexity and rapidly evolving data and algorithmic demands, strongly motivates the formulation of a prioritized research agenda to maximize their capabilities and impact. In response to this need, DOE’s Advanced Scientific Computing Research (ASCR) program in the Office of Science convened the Workshop on Basic Research Needs for Inverse Problems for Complex Systems Under Uncertainty in June 2025. This workshop brought together experts across disciplines to identify grand challenges and major opportunities in the field. Through collaborative discussions, the workshop defined transformative research directions aimed at addressing the mathematical, statistical, and computational challenges posed by inverse problems under uncertainty. As a result of these efforts, four priority research directions (PRDs) were identified to guide future research and development in this area. These PRDs, summarized below, represent a roadmap for advancing the foundational science and mathematics of inverse problems, enabling robust, scalable, and uncertainty-aware solutions that are critical for DOE applications.

97 MATHEMATICS AND COMPUTING↗

Gearbox bearing crack growth prognostics and uncertainty quantification with physics-informed machine learning

This paper introduces the extreme theory of functional connections (X-TFC), a physics-informed machine learning algorithm, and tailors it to estimate the remaining useful life (RUL) of wind turbine gearbox bearings experiencing fatigue crack growth. Unlike purely data-driven methods, X-TFC embeds a physics model, based on Head's theory in this work, into its training objective. The core of X-TFC is a random-projection single-layer neural network trained via an extreme learning machine, which requires only limited damage progression data and solves for output weights with a least-squares optimization algorithm. A composite loss function balances the network's fit to observed degradation data against the residuals of the governing crack growth differential equation, ensuring the learned damage trajectory remains physically plausible. When applied to a vibration-based health-index (HI) dataset measured during the growth of a crack on the inner ring of a high-speed bearing in a wind turbine gearbox (Bechhoefer and Dubé, 2020), X-TFC achieves near-zero prediction bias. Even when trained on only the first 10 %–20 % of the damage progression data, with sufficient physics weighting its predictions remain monotonic and smooth, delivering high prognosability and trendability. To quantify the epistemic uncertainty, we employ a Monte Carlo ensemble of independently initialized X-TFC models trained on noise-perturbed data, which yields confidence intervals around each RUL estimate and captures both model-parameter and epistemic uncertainty. In addition to a vibration-based HI, we demonstrate that the proposed framework can be directly applied to a supervisory control and data acquisition (SCADA) data-based HI (Eftekhari Milani et al., 2026) measured during similar wind turbine gearbox bearing crack faults, preserving its accuracy and interpretability. This extension shows the versatility of our approach, which is applicable to bearings of multiple gearbox manufacturers, models, and ratings using only SCADA data. By integrating domain knowledge with machine learning, X-TFC offers a rapid, reliable tool for crack prognostics. Its adaptability to other bearing failure modes, such as pitch bearing ring cracks, positions X-TFC as a powerful enabler of data-driven, physics-informed asset management in the wind energy sector and beyond.

17 WIND ENERGY↗

Thermal Conductivity Measurement of Extrusion-Printed Silver Using Modulated Photothermal Radiometry

Flexible printed electronics is a rapidly growing field with applications in conformal and flexible devices. However, the physical properties of the films created by many state-of-the-art printing methods become highly dependent on printing parameters, resulting in varying thermal properties often differing significantly from their bulk ink components. To understand the influence of the printing process, we build upon our previous work, where a noncontact optical technique, known as modulated photothermal radiometry (MPTR), was used to measure the thermal conductivity of aerosol-jet-printed thin films. In this work, we use the method to study the thermal properties of extrusion-printed silver on glass and alumina substrates. A noise-resistant data analysis fitting technique is applied using a 2-D heat transfer model. Here, the thermal conductivity measurement is validated using the Weidemann-Franz (WF) relationship from measured electrical conductivity values.

42 ENGINEERING↗

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗

Machine learning for improved current-density reconstruction from two-dimensional vector magnetic images

The reconstruction of electrical current densities from magnetic field measurements is an important technique with applications in materials science, circuit design, quality control, plasma physics, and biology. Analytic reconstruction methods exist for planar currents, but break down in the presence of high-spatial-frequency noise or large standoff distance, restricting the types of systems that can be studied. Here, we demonstrate the use of a deep convolutional neural network for current density reconstruction from two-dimensional images of vector magnetic fields acquired by a quantum diamond microscope . Trained network performance significantly exceeds analytic reconstruction for data with high noise or large standoff distances. This machine learning technique can perform quality inversions on lower-signal-to-noise-ratio data, significantly reducing the data collection time and permitting reconstructions of weaker and three-dimensional current sources. Published by the American Physical Society 2025

Reed, Niko R. (ORCID:0009000305222403)↗

Algorithms for Non-Negative Matrix Factorization on Noisy Data With Negative Values

Non-negative matrix factorization (NMF) is a dimensionality reduction technique that has shown promise for analyzing noisy data, especially astronomical data. For these datasets, the observed data may contain negative values due to noise even when the true underlying physical signal is strictly positive. Prior NMF work has not treated negative data in a statistically consistent manner, which becomes problematic for low signal-to-noise data with many negative values. In this paper we present two algorithms, Shift-NMF and Nearly-NMF, that can handle both the noisiness of the input data and also any introduced negativity. Both of these algorithms use the negative data space without clipping or masking and recover non-negative signals without any introduced positive offset that occurs when clipping or masking negative data. We demonstrate this numerically on both simple and more realistic examples, and prove that both algorithms have monotonically decreasing update rules.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DECOVALEX-2023: Task D Final Report

Task D of DECOVALEX-2023 is focused on the simulation of the coupled thermal hydraulic-mechanical (THM) behaviour in the full-scale engineered barrier system (EBS). The Horonobe EBS experiment is the demonstration of the full-scale EBS in the underground research laboratory (URL) (performed by JAEA in the Horonobe URL in Japan). Task D consisted of the three steps, a preliminary step (Step 0), simulation of the laboratory tests (Step 1) and simulation of the in-situ full-scale EBS experiment (Step 2). Since the Horonobe EBS experiment demonstrates the vertical emplacement option of the EBS, the experiment gallery is also backfilled with the backfill material. Therefore, interaction between the EBS and the backfill material can also be demonstrated, such as deformation (change of density) of the buffer material. The underground water in the Horonobe URL is saline. This fact adds chemical processes to THM behaviour. For example, mechanical properties (such as swelling pressure of the buffer material and backfill material) and hydraulic properties (such as permeability of the buffer material and backfill material) change depending on the water chemistry. Task D was therefore a challenging Task focused on not only the relatively simple THM behaviour but also complex THM behaviour including chemical processes. Six research teams (BGR, CAS, JAEA, KAERI, SNL and Taipower) participated the Task D. BGR, CAS, JAEA, KAERI and Taipower research teams selected a THM approach, while the SNL research team selected a TH approach. Step 1 involved the simulation of laboratory test results and was important to check the numerical codes developed by the research teams. Step 1 was divided into four sub steps. The simulation results through the Step 1 identified the parameters for simulation of the Step 2. Basic parameters of the materials (buffer material, backfill material, rock mass, concrete, sand) were provided by JAEA. Special parameters which research team needed were identified by back analysis of Step 1. Most notably the mechanical behaviour of swelling and displacement depended on the applied model (elastic model or elastoplastic model). Parameters such as Young’s modulus were found to need smaller values than characterised in the fundamental laboratory test results (Step 1-1, 1-2) for the elastic model. Although laboratory experiments are usually simple, test results contained some error. For example, if the saturation level is 100 % or higher, it should be considered an error. This situation was presented in the Step 1-3. A possible reason is that the buffer material is a mixture of bentonite and silica sand. When a specimen is cut to measure volume or weight, sand grains will affect the measurement data. In Step 2, boundary conditions such as temperature on the surface of the simulated overpack, heater power of the electrical heaters installed in the simulated overpack, injection pressure and inflow rate of the test water, were applied. The outer boundary conditions can be selected using measured data (injection pressure and inflow rate of the test water that is controlled by the injection systems installed in the sand layer around the buffer material and in the boundary between backfill material and concrete support). Since such measured data has some noise, research teams developed their own simplified boundary conditions. Inner boundary conditions can be selected using measured data as heater power and temperature on the surface of the simulated overpack. These data also contain some noise, so research teams developed their own simplified developed boundary conditions. Task D validated various approaches thorough the simulation of the in-situ full scale EBS system including backfill of the gallery: variations in the coupling processes (THM or THC), analysis codes, and boundary conditions. Temperature distribution in the buffer material was simulated well by all research teams. This means thermal behaviour is not sensitive to the simulation approaches. Although the water content distribution on the outside of the buffer material was well simulated by all research teams, the simulation results differ from the measured values inside the buffer material (at the centre and inside, near the simulated overpack). The buffer material is made from tap water, but in the in-situ experiment, saline groundwater infiltrates the buffer material. Therefore, the selection of the hydraulic parameters of the buffer material greatly affects the simulation results of the re saturation behaviour of the buffer material. In the Horonobe EBS experiment, measured values suitable for validating the simulation results were not obtained near the simulated overpack. When simulating the pressure and deformation of the buffer material, the measurement data is easily affected by the installation conditions of the measurement sensors, so verifying the measurement data itself remains an issue. Mechanical simulation results differ depending on whether they are considered as elastic or elastoplastic phenomena. The accuracy of measured in-situ data can be assessed by detailed analysis comparing sampling specimen analysis and measured data. The Horonobe EBS experiment is scheduled to be dismantled in the future (FY2026 and 2027). This detailed dismantling investigation will finally confirm the measured data.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Experimental demonstration of real-time electron temperature profile control in DIII-D

Future tokamak reactor operation will require the ability to maintain a given plasma scenario for extended periods of time. This will necessitate the capability to react to changes in the plasma state and return the plasma to the target scenario; the principal method to achieve this is through feedback control. Thus, it is necessary to develop and test feedback controllers for the plasma profiles that define a target scenario. In this work, a feedback controller for the electron temperature (Te) profile is tested experimentally in DIII-D. This experiment relied on the ability to ascertain the electron temperature profile in real time, which was achieved using an observer algorithm. The observer relies on both diagnostic data and a predictive model of the electron temperature profile evolution; this predictive model includes contributions from neural network surrogate models. Because of these dependencies, a number of capabilities needed to be added to the real-time PCS for DIII-D in order to support the Te profile control experiment. The neural network surrogates needed to be integrated into the PCS to be called in real time. An observer algorithm for the Te profile needed to be added and connected to the Thomson scattering system to allow access to the current state of the profile in real time. When tested, the observer was shown to produce Te profiles that are consistent with the shape of the Thomson scattering data while rejecting much of the noise in the diagnostic data. Finally, the controller itself was tested in real time. This experiment showed that the controller is capable of tracking the electron temperature target at locations across the spatial profile.

Morosohk, Shira [Oak Ridge Associated Universities↗

CovTransformer: A transformer model for SARS-CoV-2 lineage frequency forecasting

With hundreds of SARS-CoV-2 lineages circulating in the global population, there is an ongoing need for predicting and forecasting lineage frequencies and thus identifying rapidly expanding lineages. Accurate prediction would allow for more focused experimental efforts to understand pathogenicity of future dominating lineages and characterize the extent of their immune escape. Here, we first show that the inherent noise and biases in lineage frequency data make a commonly-used regression-based approach unreliable. To address this weakness, we constructed a machine learning model for SARS-CoV-2 lineage frequency forecasting, called CovTransformer, based on the transformer architecture. We designed our model to navigate challenges such as a limited amount of data with high levels of noise and bias. We first trained and tested the model using data from the UK and the USA, and then tested the generalization ability of the model to many other countries and US states. Remarkably, the trained model makes accurate predictions two months into the future with high levels of accuracy both globally (in 31 countries with high levels of sequencing effort) and at the US-state level. Our model performed substantially better than a widely used forecasting tool, the multinomial regression model implemented in Nextstrain, demonstrating its utility in SARS-CoV-2 monitoring. Assuming a newly emerged lineage is identified and assigned, our test using retrospective data shows that our model is able to identify the dominating lineages 7 weeks in advance on average before they became dominant. Overall, our work demonstrates that transformer models represent a promising approach for SARS-CoV-2 forecasting and pandemic monitoring.

60 APPLIED LIFE SCIENCES↗

Regularization via f -Divergence: An Application to Multi-Oxide Spectroscopic Analysis

In this paper, we explore the application of convolutional neural networks (CNNs) for predicting the chemical composition of complex geologic samples in a simulated Martian atmospheric environment. Specifically, we aim to characterize oxide weight percentages (wt.%) of rock samples analyzed by remote Laser-Induced Breakdown Spectroscopy (LIBS), framing the problem as a multi-target regression task . Neural networks trained on LIBS spectra are prone to overfitting due to high spectral complexity, limited labeled data, and measurement noise. While regularization is critical for improving generalization, common methods (e.g., ℓ 2 regularization) impose constraints not directly tied to data distribution properties. We propose a novel regularization method based on a specific ƒ-divergence induced by a graph-based estimator, designed to constrain the distributional discrepancy between predictions and targets. This regularizer serves a dual purpose: (a) mitigating overfitting by enforcing a constraint on the distributional difference between predictions and noisy targets, and (b) acting as an auxiliary loss that penalizes large divergences. To enable backpropagation, we develop a differentiable approximation of this particular ƒ-divergence, making the method feasible for neural networks. Experiments on ChemCam and SuperCam LIBS calibration spectra show that mathematical equation-divergence regularization outperforms or matches standard regularization methods (ℓ 1 , ℓ 2 , dropout) and the classical baseline, partial least squares (PLS). Combining ƒ-divergence regularization with standard regularization yields further performance gains, indicating that distributional regularization is useful in this context giving a promising direction for robust model training in planetary science applications. Source code is publicly available at Klein and Li (2025), https://doi.org/10.11578/dc.20250530.7.

58 GEOSCIENCES↗

Generative Thermodynamic Computing

Here, we introduce a generative modeling framework for thermodynamic computing, in which structured data are synthesized from noise by the natural time evolution of a physical system governed by Langevin dynamics. While conventional diffusion models use neural networks to perform denoising, here the information needed to generate structure from noise is encoded by the dynamics of a thermodynamic system. Training proceeds by maximizing the probability with which the computer generates the reverse of a noising trajectory, which ensures that the computer generates data with minimal heat emission. We demonstrate this framework within a digital simulation of a thermodynamic computer. If realized in analog hardware, such a system would function as a generative model that produces structured samples without the need for artificially injected noise or active control of denoising.

Whitelam, Stephen [Lawrence Berkeley National Labo↗