Search NASASearch

SEARCH · Search NASA

Results for “Data reconstruction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Real-time reconstruction of ground motion during small magnitude earthquakes: A pilot study

This study presents a pilot investigation into a novel method for reconstructing real-time ground motion during small magnitude earthquakes (M < 4.5), removing the need for computationally expensive source characterization and simulation processes to assess ground shaking. Small magnitude earthquakes, which occur frequently and can be modeled as point sources, provide ideal conditions for evaluating real-time reconstruction methods. Utilizing sparse observation data, the method applies the Gappy Auto-Encoder (Gappy AE) algorithm for efficient field data reconstruction. This is the first study to apply the Gappy AE algorithm to earthquake ground motion reconstruction. Numerical experiments conducted with SW4 simulations demonstrate the method’s accuracy and speed across varying seismic scenarios. The reconstruction performance is further validated using real seismic data from the Berkeley area in California, USA, demonstrating the potential for practical application of real-time earthquake data reconstruction using Gappy AE. As a pilot investigation, it lays the groundwork for future applications to larger and more complex seismic events.

58 GEOSCIENCES

Enhancing Sterile Neutrino Oscillation Sensitivities using SBND-PRISM at the Short-Baseline Neutrino Programme

The Short-Baseline Neutrino (SBN) Programme at Fermilab is comprised of two detectors, SBND and ICARUS, placed at $110$\,m and $600$\,m along the Booster Neutrino Beam at Fermilab. The key physics aim of the programme is to definitively test the sterile neutrino hypothesis, a proposed fourth flavour of neutrino that may explain certain experimental anomalies seen regarding standard model neutrino oscillations. To be able to detect the existence of sterile neutrinos, the uncertainties of the programme must be well constrained. To enable this, a robust analysis must be constructed that can consistently identify the correct values of systematic parameters and generate accurate predictions of what the neutrino energy spectrum at ICARUS should look like. SBND's design allows for the introduction of a technique called PRISM. In PRISM, the detector is divided into regions of different off axis angles from the beam, forming different samples where systematics impact each one in a distinct way. This allows any fits performed to obtain a better understanding of the correct value of the systematic parameters. The first analysis included in this thesis focuses on the improvements seen to the sensitivity of SBN to sterile oscillation parameters when using a PRISM configuration instead of treating SBND as a single whole. Focusing on the $5\sigma$ exclusion contour from the $\numu$ disappearance channel using three PRISM samples, an improvement on the order of $30$\,\% is seen, extending the parameter space for which the null hypothesis can be excluded. This thesis also presents a series of mock data studies comparing the abilities of SBND and PRISM analyses, concluding that in the case of simple changes between Monte Carlo (MC) and data, PRISM produces predictions of the ICARUS event rate spectrum that are more accurate and have smaller uncertainties. When moving to more realistic mock data, using different models to create the mock data than were used for the MC, the postfit predictions at ICARUS had a smaller difference between the postfit and mock data reconstructed energy spectra across all mock data samples tested. Finally, a covariance matrix defined by the maximum discrepancy between the postfit and mock data spectra at ICARUS from each of the SBND and PRISM fits across all the samples was constructed. The resultant $1\sigma$ fractional error induced by this bias systematic on the postfit spectrum has a smaller magnitude for PRISM than SBND, with the improvements ranging from $2.21$\,\% to $6.26$\,\% depending on the mock data samples. This summarises the reduction in systematic error when using PRISM instead of treating SBND as a single detector. *************************************** AUTHOR = Slater, Bethany University of Liverpool b.slater2@liverpool.ac.uk TITLE = Enhancing Sterile Neutrino Oscillation Sensitivities using SBND-PRISM at the Short-Baseline Neutrino Programme PAGES = 218 NOTE = Ph.D. University of Liverpool March 2026 ABSTRACT = The Short-Baseline Neutrino (SBN) Programme at Fermilab is comprised of two detectors, SBND and ICARUS, placed at $110$\,m and $600$\,m along the Booster Neutrino Beam at Fermilab. The key physics aim of the programme is to definitively test the sterile neutrino hypothesis, a proposed fourth flavour of neutrino that may explain certain experimental anomalies seen regarding standard model neutrino oscillations. To be able to detect the existence of sterile neutrinos, the uncertainties of the programme must be well constrained. To enable this, a robust analysis must be constructed that can consistently identify the correct values of systematic parameters and generate accurate predictions of what the neutrino energy spectrum at ICARUS should look like. SBND's design allows for the introduction of a technique called PRISM. In PRISM, the detector is divided into regions of different off axis angles from the beam, forming different samples where systematics impact each one in a distinct way. This allows any fits performed to obtain a better understanding of the correct value of the systematic parameters. The first analysis included in this thesis focuses on the improvements seen to the sensitivity of SBN to sterile oscillation parameters when using a PRISM configuration instead of treating SBND as a single whole. Focusing on the $5\sigma$ exclusion contour from the $\numu$ disappearance channel using three PRISM samples, an improvement on the order of $30$\,\% is seen, extending the parameter space for which the null hypothesis can be excluded. This thesis also presents a series of mock data studies comparing the abilities of SBND and PRISM analyses, concluding that in the case of simple changes between Monte Carlo (MC) and data, PRISM produces predictions of the ICARUS event rate spectrum that are more accurate and have smaller uncertainties. When moving to more realistic mock data, using different models to create the mock data than were used for the MC, the postfit predictions at ICARUS had a smaller difference between the postfit and mock data reconstructed energy spectra across all mock data samples tested. Finally, a covariance matrix defined by the maximum discrepancy between the postfit and mock data spectra at ICARUS from each of the SBND and PRISM fits across all the samples was constructed. The resultant $1\sigma$ fractional error induced by this bias systematic on the postfit spectrum has a smaller magnitude for PRISM than SBND, with the improvements ranging from $2.21$\,\% to $6.26$\,\% depending on the mock data samples. This summarises the reduction in systematic error when using PRISM instead of treating SBND as a single detector.

Slater, Bethany [Liverpool U.]

Enhancing Sterile Neutrino Oscillation Sensitivities using SBND-PRISM at the Short-Baseline Neutrino Programme

The Short-Baseline Neutrino (SBN) Programme at Fermilab is comprised of two detectors, SBND and ICARUS, placed at $110$\,m and $600$\,m along the Booster Neutrino Beam at Fermilab. The key physics aim of the programme is to definitively test the sterile neutrino hypothesis, a proposed fourth flavour of neutrino that may explain certain experimental anomalies seen regarding standard model neutrino oscillations. To be able to detect the existence of sterile neutrinos, the uncertainties of the programme must be well constrained. To enable this, a robust analysis must be constructed that can consistently identify the correct values of systematic parameters and generate accurate predictions of what the neutrino energy spectrum at ICARUS should look like. SBND's design allows for the introduction of a technique called PRISM. In PRISM, the detector is divided into regions of different off axis angles from the beam, forming different samples where systematics impact each one in a distinct way. This allows any fits performed to obtain a better understanding of the correct value of the systematic parameters. The first analysis included in this thesis focuses on the improvements seen to the sensitivity of SBN to sterile oscillation parameters when using a PRISM configuration instead of treating SBND as a single whole. Focusing on the $5\sigma$ exclusion contour from the $\numu$ disappearance channel using three PRISM samples, an improvement on the order of $30$\,\% is seen, extending the parameter space for which the null hypothesis can be excluded. This thesis also presents a series of mock data studies comparing the abilities of SBND and PRISM analyses, concluding that in the case of simple changes between Monte Carlo (MC) and data, PRISM produces predictions of the ICARUS event rate spectrum that are more accurate and have smaller uncertainties. When moving to more realistic mock data, using different models to create the mock data than were used for the MC, the postfit predictions at ICARUS had a smaller difference between the postfit and mock data reconstructed energy spectra across all mock data samples tested. Finally, a covariance matrix defined by the maximum discrepancy between the postfit and mock data spectra at ICARUS from each of the SBND and PRISM fits across all the samples was constructed. The resultant $1\sigma$ fractional error induced by this bias systematic on the postfit spectrum has a smaller magnitude for PRISM than SBND, with the improvements ranging from $2.21$\,\% to $6.26$\,\% depending on the mock data samples. This summarises the reduction in systematic error when using PRISM instead of treating SBND as a single detector. *************************************** AUTHOR = Slater, Bethany University of Liverpool b.slater2@liverpool.ac.uk TITLE = Enhancing Sterile Neutrino Oscillation Sensitivities using SBND-PRISM at the Short-Baseline Neutrino Programme PAGES = 218 NOTE = Ph.D. University of Liverpool March 2026 ABSTRACT = The Short-Baseline Neutrino (SBN) Programme at Fermilab is comprised of two detectors, SBND and ICARUS, placed at $110$\,m and $600$\,m along the Booster Neutrino Beam at Fermilab. The key physics aim of the programme is to definitively test the sterile neutrino hypothesis, a proposed fourth flavour of neutrino that may explain certain experimental anomalies seen regarding standard model neutrino oscillations. To be able to detect the existence of sterile neutrinos, the uncertainties of the programme must be well constrained. To enable this, a robust analysis must be constructed that can consistently identify the correct values of systematic parameters and generate accurate predictions of what the neutrino energy spectrum at ICARUS should look like. SBND's design allows for the introduction of a technique called PRISM. In PRISM, the detector is divided into regions of different off axis angles from the beam, forming different samples where systematics impact each one in a distinct way. This allows any fits performed to obtain a better understanding of the correct value of the systematic parameters. The first analysis included in this thesis focuses on the improvements seen to the sensitivity of SBN to sterile oscillation parameters when using a PRISM configuration instead of treating SBND as a single whole. Focusing on the $5\sigma$ exclusion contour from the $\numu$ disappearance channel using three PRISM samples, an improvement on the order of $30$\,\% is seen, extending the parameter space for which the null hypothesis can be excluded. This thesis also presents a series of mock data studies comparing the abilities of SBND and PRISM analyses, concluding that in the case of simple changes between Monte Carlo (MC) and data, PRISM produces predictions of the ICARUS event rate spectrum that are more accurate and have smaller uncertainties. When moving to more realistic mock data, using different models to create the mock data than were used for the MC, the postfit predictions at ICARUS had a smaller difference between the postfit and mock data reconstructed energy spectra across all mock data samples tested. Finally, a covariance matrix defined by the maximum discrepancy between the postfit and mock data spectra at ICARUS from each of the SBND and PRISM fits across all the samples was constructed. The resultant $1\sigma$ fractional error induced by this bias systematic on the postfit spectrum has a smaller magnitude for PRISM than SBND, with the improvements ranging from $2.21$\,\% to $6.26$\,\% depending on the mock data samples. This summarises the reduction in systematic error when using PRISM instead of treating SBND as a single detector.

Slater, Bethany [Liverpool U.]

Enhancing Sterile Neutrino Oscillation Sensitivities using SBND-PRISM at the Short-Baseline Neutrino Programme

The Short-Baseline Neutrino (SBN) Programme at Fermilab is comprised of two detectors, SBND and ICARUS, placed at $110$\,m and $600$\,m along the Booster Neutrino Beam at Fermilab. The key physics aim of the programme is to definitively test the sterile neutrino hypothesis, a proposed fourth flavour of neutrino that may explain certain experimental anomalies seen regarding standard model neutrino oscillations. To be able to detect the existence of sterile neutrinos, the uncertainties of the programme must be well constrained. To enable this, a robust analysis must be constructed that can consistently identify the correct values of systematic parameters and generate accurate predictions of what the neutrino energy spectrum at ICARUS should look like. SBND's design allows for the introduction of a technique called PRISM. In PRISM, the detector is divided into regions of different off axis angles from the beam, forming different samples where systematics impact each one in a distinct way. This allows any fits performed to obtain a better understanding of the correct value of the systematic parameters. The first analysis included in this thesis focuses on the improvements seen to the sensitivity of SBN to sterile oscillation parameters when using a PRISM configuration instead of treating SBND as a single whole. Focusing on the $5\sigma$ exclusion contour from the $\numu$ disappearance channel using three PRISM samples, an improvement on the order of $30$\,\% is seen, extending the parameter space for which the null hypothesis can be excluded. This thesis also presents a series of mock data studies comparing the abilities of SBND and PRISM analyses, concluding that in the case of simple changes between Monte Carlo (MC) and data, PRISM produces predictions of the ICARUS event rate spectrum that are more accurate and have smaller uncertainties. When moving to more realistic mock data, using different models to create the mock data than were used for the MC, the postfit predictions at ICARUS had a smaller difference between the postfit and mock data reconstructed energy spectra across all mock data samples tested. Finally, a covariance matrix defined by the maximum discrepancy between the postfit and mock data spectra at ICARUS from each of the SBND and PRISM fits across all the samples was constructed. The resultant $1\sigma$ fractional error induced by this bias systematic on the postfit spectrum has a smaller magnitude for PRISM than SBND, with the improvements ranging from $2.21$\,\% to $6.26$\,\% depending on the mock data samples. This summarises the reduction in systematic error when using PRISM instead of treating SBND as a single detector. *************************************** AUTHOR = Slater, Bethany University of Liverpool b.slater2@liverpool.ac.uk TITLE = Enhancing Sterile Neutrino Oscillation Sensitivities using SBND-PRISM at the Short-Baseline Neutrino Programme PAGES = 218 NOTE = Ph.D. University of Liverpool March 2026 ABSTRACT = The Short-Baseline Neutrino (SBN) Programme at Fermilab is comprised of two detectors, SBND and ICARUS, placed at $110$\,m and $600$\,m along the Booster Neutrino Beam at Fermilab. The key physics aim of the programme is to definitively test the sterile neutrino hypothesis, a proposed fourth flavour of neutrino that may explain certain experimental anomalies seen regarding standard model neutrino oscillations. To be able to detect the existence of sterile neutrinos, the uncertainties of the programme must be well constrained. To enable this, a robust analysis must be constructed that can consistently identify the correct values of systematic parameters and generate accurate predictions of what the neutrino energy spectrum at ICARUS should look like. SBND's design allows for the introduction of a technique called PRISM. In PRISM, the detector is divided into regions of different off axis angles from the beam, forming different samples where systematics impact each one in a distinct way. This allows any fits performed to obtain a better understanding of the correct value of the systematic parameters. The first analysis included in this thesis focuses on the improvements seen to the sensitivity of SBN to sterile oscillation parameters when using a PRISM configuration instead of treating SBND as a single whole. Focusing on the $5\sigma$ exclusion contour from the $\numu$ disappearance channel using three PRISM samples, an improvement on the order of $30$\,\% is seen, extending the parameter space for which the null hypothesis can be excluded. This thesis also presents a series of mock data studies comparing the abilities of SBND and PRISM analyses, concluding that in the case of simple changes between Monte Carlo (MC) and data, PRISM produces predictions of the ICARUS event rate spectrum that are more accurate and have smaller uncertainties. When moving to more realistic mock data, using different models to create the mock data than were used for the MC, the postfit predictions at ICARUS had a smaller difference between the postfit and mock data reconstructed energy spectra across all mock data samples tested. Finally, a covariance matrix defined by the maximum discrepancy between the postfit and mock data spectra at ICARUS from each of the SBND and PRISM fits across all the samples was constructed. The resultant $1\sigma$ fractional error induced by this bias systematic on the postfit spectrum has a smaller magnitude for PRISM than SBND, with the improvements ranging from $2.21$\,\% to $6.26$\,\% depending on the mock data samples. This summarises the reduction in systematic error when using PRISM instead of treating SBND as a single detector.

Slater, Bethany [Liverpool U.]

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES

WRF-Chem & Data Assimilation: Reconstructing a Historic Dispersion Event

CONCLUSIONS • For this particular case and use, 3D-VAR was the most useful data assimilation method applied for modeling this plume. • Slight downside was the plume was modeled to move too quickly. • Whether this result is generalizable may be debatable based on the circumstances. • The atmospheric state altered how effective the data assimilation algorithms performed. • Nudging and 3D-VAR are not mutually exclusive, but their errors for wind speed may compound

Thomas, Andrew M. [Savannah River National Laborat

Securing Federated Learning Against Active Reconstruction Attacks

Federated Learning (FL) has amassed notable attention for its ability to preserve user privacy while emphasizing the retainment of model training efficiency. Due to this potential, FL has been integrated in many domains, such as healthcare, finance, law, and industrial engineering, where data cannot be easily exchanged due to sensitive information and strict privacy laws. However, current research has indicated that FL protocols are easily compromised by active data reconstruction attacks employed by actively dishonest servers. The malicious modification of global model parameters allows an actively dishonest server to obtain a direct copy of users’ private data via gradient inversion. Here, this class of attacks is highly underexplored and continues to be a major challenge due to the intense threat model. In this paper, we propose OASIS as a scalable and modality-agnostic defense based on data augmentation that counteracts active data reconstruction attacks while preserving model performance. To generalize our defense, we uncover the intuition behind gradient inversion that enables these attacks and theoretically establish the conditions by which the defense can be considered robust regardless of attack design. From this, we formulate our defense with data augmentation that illustrates its ability to undermine the attack principle. We evaluate OASIS on five real-world datasets–two image-based (ImageNet and CIFAR100) and three text-based (Wikitext, Stack Overflow, and Shakespeare)–which span diverse uses cases such as vision tasks and language modeling. Comprehensive evaluations on these datasets exhibit the efficacy of OASIS and highlight its feasibility as a solution.

97 MATHEMATICS AND COMPUTING

Integrated top-down process and voxel-based microstructure modeling for Ti-6Al-4V in laser wire direct energy deposition process

Laser-wire metal additive manufacturing (AM) is one of the ideal direct energy deposition (DED) processes for creating large-scale parts with a medium level of complexity. However, the DED process involves complex thermal signatures and wide length scales making the fabrication of realistic AM components and part qualification often reliant on experimental trial-and-error optimization. While experimental measurements over the full volume of a part are valuable and necessary, measuring the entire area of a part is significantly laborious and practically infeasible, particularly for large parts in terms of cost and rapid qualification. Therefore, in this work, we developed an effective thermal and microstructure modeling framework based on the Johnson–Mehl-Avrami-Kolmogorov (JMAK) and Koistinen & Marburger (KM) models through a top-down approach that considers plate distortion-affected thermal profiles. A voxel-by-voxel simulation method is used to predict individual phase fractions of Ti-6Al-4 V. The predicted results were validated through detailed metallurgical measurements. A combined voxel-by-voxel approach with a sparse data reconstruction technique produced a near-perfect reconstruction of the original data. This approach anticipates a significant reduction in data points and computation time and resources. Lastly, we conclude with potential extensions of this work to other modeling efforts.

36 MATERIALS SCIENCE

Fast Hyperspectral Neutron Tomography

Hyperspectral neutron computed tomography is a tomographic imaging technique in which thousands of wavelength-specific neutron radiographs are measured for each tomographic view. In conventional hyperspectral reconstruction, data from each neutron wavelength bin are reconstructed separately, which is extremely time-consuming. These reconstructions often suffer from poor quality due to low signal-to-noise ratios. Consequently, material decomposition based on these reconstructions tends to produce inaccurate estimates of the material spectra and erroneous volumetric material separation. In this paper, we present two novel algorithms for processing hyperspectral neutron data: fast hyperspectral reconstruction and fast material decomposition. Both algorithms rely on a subspace decomposition procedure that transforms hyperspectral views into low-dimensional projection views within an intermediate subspace, where tomographic reconstruction is performed. The use of subspace decomposition dramatically reduces reconstruction time while reducing both noise and reconstruction artifacts. We apply our algorithms to both simulated and measured neutron data and demonstrate that they reduce computation and improve the quality of the results relative to conventional methods.

Chowdhury, Mohammad Samin Nur [Purdue University]

A Survey on Error-Bounded Lossy Compression for Scientific Datasets

Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.

Error-Bounded Lossy Compression

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder

Get Non-Real: Randomized Sketching for High-Dimensional Non-Real Valued Data (Final Report)

In our final report for DE-C0022186, we describe the work we did on this grant towards the goals we proposed. Our first goal was characterizing fundamental limits for sketching of discrete high-dimensional matrices with low-dimensional structures. Our second main goal was designing algorithms for data reconstruction from sketches. We focus on approaches that are either specifically designed for non-real-valued data (binary, finite field) or that will translate more readily to that setting.

97 MATHEMATICS AND COMPUTING

Synthetic modeling of soft x-ray emissivity for magnetic island analysis in LTX-β

We present a synthetic soft x-ray (SXR) forward-modeling framework that characterizes the emissivity structure of rotating magnetic islands in the Lithium Tokamak eXperiment-β across space and time. Magnetic islands associated with tearing modes produce modulations in line-integrated SXR brightness. Traditional tomographic methods struggle to resolve these structures in devices with limited sightlines, resulting in an under-determined inversion problem, or otherwise necessitate equilibrium reconstruction data, preventing their use in active control. Our approach removes this challenge by combining a fully three-dimensional ray-tracing model with a time-dependent emissivity prescription of a $m/n$ = 2/1 magnetic island. The model is validated through helical island geometry, rotation measured by magnetic diagnostics, and equilibrium constraints from PSI-Tri reconstructions. Synthetic brightness signals are generated for each photodiode sightline from a modeled emissivity profile and directly compared with experimental data from a tangential midplane SXR array. By fitting the synthetic diagnostic output to the observed brightness across time, we simultaneously infer key geometric parameters-such as island radial location, island width, and rotation frequency-without requiring full tomographic reconstruction, all with an average absolute deviation of less than 3%. This work demonstrates that forward modeling with a single tangential array can extract key island parameters in a spherical tokamak, providing a computationally efficient alternative to conventional SXR tomography and providing a pathway toward real-time magnetic-island characterization in future devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Online randomized interpolative decomposition with a posteriori error estimator for temporal PDE data reduction

Traditional low-rank approximation is a powerful tool for compressing large data matrices that arise in simulations of partial differential equations (PDEs), but suffers from high computational cost and requires several passes over the PDE data. The compressed data may also lack interpretability thus making it difficult to identify feature patterns from the original data. Here, to address these issues, we present an online randomized algorithm to compute the interpolative decomposition (ID) of large-scale data matrices in situ. Compared to previous randomized IDs that used the QR decomposition to determine the column basis, we adopt a streaming ridge leverage score-based column subset selection algorithm that dynamically selects proper basis columns from the data and thus avoids an extra pass over the data to compute the coefficient matrix of the ID. In particular, we adopt a single-pass error estimator based on the non-adaptive Hutch++ algorithm to provide real-time error approximation for determining the best coefficients. As a result, our approach only needs a single pass over the original data and thus is suitable for large and high-dimensional matrices stored outside of core memory or generated in PDE simulations. A strategy to improve the accuracy of the reconstructed data gradient, when desired, within the ID framework is also presented. We provide numerical experiments on turbulent channel flow and ignition simulations, and on the NSTX Gas Puff Image dataset, comparing our algorithm with the offline ID algorithm to demonstrate its utility in real-world applications.

Column subset selection

Abundance and properties of dark radiation from the cosmic microwave background

We study the cosmological signatures of new light relics that are collisionless like standard neutrinos or are strongly interacting. We provide a simple and succinct rephrasing of their physical effects in the cosmic microwave background, as well as the resulting parameter degeneracies with other cosmological parameters, in terms of the total radiation abundance and the fraction thereof that freely streams. In these more general terms, interacting and noninteracting light relics are differentiated by their respective decrease and increase of the free-streaming fraction, and, moreover, the scale-dependent interplay thereof with a common, correlated reduction of the fraction of matter in baryons. We then derive updated constraints on various dark-radiation scenarios with the latest cosmological observations, employing this language to identify the physical origin of the impact of each dataset. The “PR4” reanalyses of Planck CMB data prefer a larger primordial helium yield and therefore also slightly more radiation than the 2018 analysis; we investigate the differences between the two releases that drives these shifts. Smaller free-streaming fractions are disfavored by the excess lensing of the CMB measured in lensing reconstruction data from Planck and the Atacama Cosmology Telescope. On the other hand, baryon acoustic oscillation measurements from the Dark Energy Spectroscopic Instrument drive marginal detections of new, strongly interacting light relics due to that data's preference for lower matter fractions. Finally, we forecast measurements from the CMB-S4 experiment.

cosmological parameters from CMBR

Search for 2p2h Interactions in the NOνA Near Detector

The physics of 2p2h interactions and their contribution to the NO$\nu$A near detector data are not fully understood. This study attempts to shed some light on these interactions and the accuracy of the models used to simulate them through a search for a specific 2p2h interaction in the NO$\nu$A near detector. By performing an event selection algorithm based on particle identifier algorithms run over reconstructed data, a signal region is created to minimize the background while maximizing the number of 2p2h events where a muon neutrino interacts with a neutron and a proton coupled by a meson exchange current and produces two protons and one muon. In the signal region, separation is found between the signal events and the background in plots of the angles between the protons and the muon. Although a full statistical analysis is not completed in this study, comparing the angle plots for simulation and data shows that the model used to simulate the events reasonably approximates reality and that the near detector data likely includes signal events. Signal events are also identified in event displays, further indicating that there is some contribution of the signal to the overall NO$\nu$A near detector data.

Gable, Kyle