Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

The 4D Camera: An 87 kHz Direct Electron Detector for Scanning/Transmission Electron Microscopy

We describe the development, operation, and application of the 4D Camera—a 576 by 576 pixel active pixel sensor for scanning/transmission electron microscopy which operates at 87,000 Hz. The detector generates data at ~480 Gbit/s which is captured by dedicated receiver computers with a parallelized software infrastructure that has been implemented to process the resulting 10–700 Gigabyte-sized raw datasets. The back illuminated detector provides the ability to detect single electron events at accelerating voltages from 30 to 300 kV. Through electron counting, the resulting sparse data sets are reduced in size by 10--300× compared to the raw data, and open-source sparsity-based processing algorithms offer rapid data analysis. The high frame rate allows for large and complex scanning diffraction experiments to be accomplished with typical scanning transmission electron microscopy scanning parameters.

47 OTHER INSTRUMENTATION↗

BoBa

BoBa is a C++ software library for working with large matrices, tensors, and tensor decompositions. The library provides tools for dense matrix and tensor operations, tensor decompositions, and tensor decomposition methods that support modern CPU and GPU architectures. It includes portable abstractions for linear algebra, tensor algebra, and multidimensional computation. BoBa is intended for scientific computing applications that involve large multidimensional data sets or high dimensional mathematical models. Its capabilities support tasks such as data compression, linear algebra, efficient numerical computation, and the development of scalable algorithms for heterogeneous hardware. Tutorials, tests, and example applications are included to help users learn and apply the library.

Yao, Jin [Lawrence Livermore National Laboratory (↗

The Foundational Industrial Energy Dataset (FIED): Open-Source Data on Industrial Facilities

The state of data on industrial energy use has co-evolved over several decades with the demands of industrial energy analysis. The most recent development - analysis in support of decarbonizing the industrial sector - has changed the characteristics of industrial data that are useful for analysts and model developers. Although data and its collection processes may be cast from a conventional viewpoint as objective and free from the influence of social dynamics, this provides an incomplete picture of not only the processes by which information is generated, but also the limitations and opportunities of data to be useful for analysis. The foundational industry energy data set (FIED) is a result of the confluence of trends in open data and the demand for higher resolution industrial energy analysis. The general approach to compiling the FIED involves accessing, filtering, and formatting data published by federal organizations on the Internet for public use. Unlike most industrial energy datasets, which are published by the U.S. Energy Information Administration (EIA), the FIED relies on core datasets from the U.S. Environmental Protection Agency (EPA). The FIED addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization - including estimates of energy use, greenhouse gas emissions, and design capacities - for facilities that are identified by latitude and longitude. This enables local-level analysis of existing combustion equipment, as well as regional comparisons with traditional industrial energy data estimates. The report summarizes the general logic behind compiling the FIED. The FIED itself and its Python code are available from OpenEI and GitHub, respectively.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Computed Tomography Scanning and Geophysical Measurements of the Integrated Mid-Continent Stacked Carbon Storage Hub Sleepy Hollow Reagan Unit 86A Well

The Computed Tomography (CT) facilities, the Multi-Sensor Core Logger (MSCL), and the Geologic Storage Core Flow laboratory at the National Energy Technology Laboratory (NETL) in Morgantown, West Virginia and Pittsburgh, Pennsylvania were used to characterize a core through Upper Pennsylvanian strata (limestones, mudstones, and sandstones) from the Sleepy Hollow Reagan Unit (SHRU) 86A well in southwest-central Nebraska. The Integrated Mid-Continent Stack Carbon Storage Hub (IMSCS-HUB) core from the vertical well was obtained as part of the Department of Energy’s (DOE) effort to assess the feasibility of stacked storage complexes in Nebraska and Kansas to support a commercial-scale CO2 storage hub. The Sleepy Hollow Field (SHF) is one of three sites within the IMSCS-HUB corridor. Bulk scans of core were obtained from the IMSCS-HUB SHRU 86A well. This report, and the associated scans, include detailed datasets not typically made available to the public. The data sets presented in this report can be accessed from NETL's Energy Data eXchange (EDX) online system.

58 GEOSCIENCES↗

Geothermal Play Fairway Analysis of Low-Temperature Resources for Sedimentary Basin Geothermal Play Types: An Example in the Denver Basin

This project is part of a nationwide effort to highlight the advantages of incorporating low-temperature geothermal resource evaluation into the implementation of combined heat and power (CHP), and geothermal direct use (GDU) technologies (e.g., space heating and/or cooling). The initiative aims to hasten the nation's decarbonization process by exploring the potential for using low-temperature geothermal resources (< 150 Degrees Celsius) in selected sedimentary basins that have several population centers. The Play Fairway Analysis (PFA) techniques were modified from earlier studies of sedimentary basin geothermal play types (SBGPTs) that assessed the viability of low-temperature resources. The decision-making process for leveraging low-temperature geothermal resources for GDU and CHP applications is complex and considers a variety of factors, including geological, economic, and risk criteria. This study covers workflows, relevant datasets, python code, and both common and composite maps used to create low-temperature geothermal resource favorability maps for the Denver Basin, which extends across Colorado, Nebraska, and Wyoming. The replication of these methodologies in other SBGPTs can evaluate potential for low-temperature resources. The proposed geothermal PFA approach for low-temperature geothermal resources includes: (1) identifying available relevant data and grouping data sets into PFA criteria (e.g., geological, economic, and risk criteria); (2) analyzing data gaps enable future focalized exploration; (3) performing uncertainty quantification; (4) weighting relevant data; (5) developing favorability and common risk maps for low-temperature geothermal resources to identify potential locations for more focused data collection. This project will facilitate future deployment of CHP and GDU by providing data, tools, and a workflow applicable to low-temperature geothermal resources in sedimentary basins.

15 GEOTHERMAL ENERGY↗

Validation of HyRAM+ Version 5.1 Physics Models

The Hydrogen Plus Other Alternative Fuels Risk Assessment Models (HyRAM+) software has seen various improvements and additional physics capabilities since validation against experimental data was last published for HyRAM v3.1. Notably, HyRAM+ now includes four models allowing for the calculation of overpressure resulting from vapor cloud explosions from unconfined jet releases. As with the previous HyRAM validation report, validation data was gathered from available published literature and tested against HyRAM+ capabilities. The validation comparisons include tank blowdown, unignited dispersion jet plume, ignited jet flame, and enclosed accumulation and overpressure. The unconfined overpressure calculations in HyRAM+ v5.1.1 generally show good agreement with many of the experimental data sets for all four unconfined overpressure models, though HyRAM+ overpredicts the experimental data for small and cryogenic hydrogen releases. The comparisons for the other HyRAM+ physics models are largely unchanged from the previously published validation report.

08 HYDROGEN↗

Orbital-Radar v1.0.0: a tool to transform suborbital radar observations to synthetic EarthCARE cloud radar data

The Earth Cloud, Aerosol and Radiation Explorer (EarthCARE) satellite developed by the European Space Agency (ESA) and the Japan Aerospace Exploration Agency (JAXA) launched in May 2024 carries a novel 94 GHz cloud profiling radar (CPR) with Doppler capability. This work describes the open-source instrument simulator Orbital-Radar, which transforms high-resolution radar data from field observations or forward simulations of numerical models to CPR primary measurements and uncertainties. The transformation accounts for sampling geometry and surface effects. We demonstrate Orbital-Radar's ability to provide realistic CPR views of typical cloud and precipitation scenes. The presented case studies show small-scale convection, marine stratus clouds, and Arctic mixed-phase cloud cases. These results provide valuable insights into the capabilities and challenges of the EarthCARE CPR mission and its advantages over the CloudSat CPR. Finally, Orbital-Radar allows for evaluating kilometre-scale numerical weather prediction models with EarthCARE CPR observations. So, Orbital-Radar can generate calibration and validation (Cal/Val) data sets already pre-launch. Nevertheless, an evaluation of synthetic CPR output data to accurate EarthCARE CPR data is missing.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Eliminates Reanalysis Warm Bias and Reveals Weaker Winter Surface Cooling Over Arctic Sea Ice

The surface energy budget governs Arctic sea-ice growth/melt, yet observations are sparse, and reanalysis data sets suffer from systematic biases. Here, we train a neural network with observational data to bias-correct hourly ERA5 fluxes over Arctic ice-covered regions (≥70°N; sea-ice concentration >80%) for 1994–2024. Training data cover two full seasonal cycles and different sea-ice regimes. The neural network reduces RMSE for net shortwave radiation by ∼40%, downward longwave radiation by ∼16% and the total surface energy budget by ∼55%, eliminating the wintertime warm bias of ∼4 K in ERA5. Wintertime surface cooling is reduced by ∼50%, yielding thermodynamic ice-growth estimates of ∼80–120 cm, consistent with SMOS–CryoSat satellite thickness increases and in contrast to the 150–200 cm growth implied by ERA5. Our bias-corrected data capture the observed clear/cloudy states of the winter boundary layer and can be used to study Arctic climatology, evaluate climate models and drive sea-ice-ocean models.

Hossain, Akil [Alfred Wegener Institute for Polar ↗

Identification of Distorted Gamma-Ray Signature Patterns Using Digital Filtering and Auto-Associative Memory Implemented with a Hopfield Neural Network

The detection and identification of radioactive sources in search applications involve analyzing passive gamma-ray emissions from high-level radioactive materials. This process uses a mobile detector-spectrometer in a complex field test environment. Recently, the use of artificial intelligence for gamma-ray spectrum analysis has shown promising results. However, challenges persist in identifying isotopic signatures from spectral measurements that may be distorted due to source shielding, random variations in natural radioactive background, or insufficient measurement time to obtain clear spectral lines. Here, this paper presents a novel intelligent signature recognition method that combines digital filtering techniques with an artificial Hopfield Neural Network (HNN). The HNN leverages auto-associative memory to store training sample patterns and match them with incoming gamma spectra from distorted sources. It restores the testing sources’ measurements by finding the closest matching signature patterns in the spectral library. Before HNN recognition, the measured spectrum undergoes preprocessing with a digital image filter to reduce fluctuations. Performance of the proposed method is evaluated using a set of gamma-ray spectra measured with a sodium iodide detector. The data collected include measurements from six pure samples: 241 Am, 60 Co, 137 Cs, 192 Ir, 239 Pu, and 235 U, which are used for training and validation (i.e. six cases). Additionally, the data set contains 24 distorted synthesized sources with various fluctuating backgrounds. Test results demonstrate the potential of the proposed method to accurately recognize the correct isotope with high precision, achieving an accuracy rate exceeding 85%. Furthermore, the proposed method exhibits superior performance compared to the conventional multiple regression fitting and simple feedforward neural network methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

In real-world applications like spectrum management and interference detection, dealing with unseen electromagnetic waveforms is critical. Although some methods attempt to simulate open set data using generator models, they face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. This results in difficulties capturing distinctive features across classes, especially in dynamic scenarios where new classes emerge. To detect unseen waveforms, we propose combining time and frequency domain features with cosine similarity loss to enhance feature distinctiveness and enabling more accurate predictions. This approach efficiently captures more comprehensive information than single-domain representations or approaches without cosine loss. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10\% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Harmonic analysis of discrete tracers of large-scale structure

It is commonplace in cosmology to analyze fields projected onto the celestial sphere, and in particular density fields that are defined by a set of points e.g. galaxies. When performing an harmonic-space analysis of such data (e.g. an angular power spectrum) using a pixelized map one has to deal with aliasing of small-scale power and pixel window functions. We compare and contrast the approaches to this problem taken in the cosmic microwave background and large-scale structure communities, and advocate for a direct approach that avoids pixelization. We describe a method for performing a pseudo-spectrum analysis of a galaxy data set and show that it can be implemented efficiently using well-known algorithms for special functions that are suited to acceleration by graphics processing units (GPUs). The method returns the same spectra as the more traditional map-based approach if in the latter the number of pixels is taken to be sufficiently large and the mask is well sampled. The method is readily generalizable to cross-spectra and higher-order functions. It also provides a convenient route for distributing the information in a galaxy catalog directly in harmonic space, as a complement to releasing the configuration-space positions and weights, and a route to spectral apodization. Finally, we make public a code enabling the application of our method to existing and upcoming datasets.

79 ASTRONOMY AND ASTROPHYSICS↗

G-Band Radar Demonstration for Microphysics Field Campaign Report

The G-Band Radar Demonstration for Microphysics (GRDM) campaign took place at the Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) from March 15 to April 30, 2024. This was a deployment of two of NASA’s Jet Propulsion Laboratory (JPL) radars and one radar from Brookhaven National Laboratory to demonstrate the utility of high-frequency millimeter-wave radars for remote sensing of stratocumulus microphysical properties. The radars were deployed on the Ellen Browning Scripps Memorial Pier alongside the AMF instruments (Figure 1). The radars include a Ka-band (35 GHz), W-band (94 GHz), and four G-band (158, 165, 174, and 240 GHz) channels. The 240 GHz and W-band channels provide complete Doppler spectra, which are useful for advanced analysis. The 158-175 GHz channels are sensitive to the water vapor profile and are useful for attenuation correction. These radars complement the high-sensitivity ARM KAZR. The goal of the deployment was to observe drizzling stratocumulus and demonstrate the capabilities of the multifrequency radar data set to constrain profiles of liquid water content and drizzle drop characteristic size. The data are still being analyzed. The methodology to derive the cloud and precipitation parameters will exploit differential attenuation and differential reflectivity between low-frequency (Ka-band) and high-frequency (G-band) channels. The method will also exploit the capability of the G-band observations to constrain the attenuation due to water vapor. These observations will quantify the capabilities and limitations of the emerging technology of G-band radars for constraint stratocumulus cloud microphysics, which are key to constraining aerosol-cloud-precipitation interactions and low-cloud climate feedback.

47 OTHER INSTRUMENTATION↗

ACE-ENA: Fast Liquid Water Content

These data were collected during the Aerosol and Cloud Experiments in the Eastern North Atlantic field campaign as part of ARM Aerial Facility deployment (ACE-ENA, https://www.arm.gov/research/campaigns/aaf2017ace-ena). The ARM Aerial Facility Gulfstream-1 was deployed at Lajes Air Base (IATA: TER, ICAO: LPLA), on Terceira Island in the Azores, Portugal, for the two Intensive Observation Periods from June 20 through July 22, 2017 (IOP#1) and from January 11 through February 22, 2018 (IOP#2). The G-1 aircraft performed 20+19 research flights over the ARM Eastern North Atlantic (ENA) site and Atlantic Ocean to measure atmospheric turbulence, cloud water content and drop size distributions, aerosol precursor gases, aerosol chemical composition and size distributions. The current data set presents re-processed Particle Volume Monitor PVM-100A (aka Gerber probe) data: Liquid Water Content (LWC), Particle Surface Area (PSA), and the effective droplet radius (re) averaged to 50 Hz, 10 Hz, and 1 Hz.

54 ENVIRONMENTAL SCIENCES↗

CACTI: Fast Liquid Water Content

These data were collected during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI; https://www.arm.gov/research/campaigns/amf2018cacti ) field campaign in the Sierras de Córdoba mountain range of north-central Argentina as part of the ARM Aerial Facility (AAF) deployment. The ARM Aerial Facility Gulfstream-1 was operated from Las Higueras Airport (IATA: RCU, ICAO: SAOC), Río Cuarto, Córdoba, Argentina, for the Intensive Observation Period (IOP) from Nov. 1 through Dec. 15, 2018. The G-1 aircraft performed 22 research flights over the first ARM Mobile Facility (AMF1) location in the Sierras de Córdoba mountain range to measure atmospheric state and turbulence, cloud water content and droplet size distributions, aerosol precursor gases, and aerosol chemical composition and size distributions. The current data set presents re-processed Particle Volume Monitor PVM-100A (aka Gerber probe) data: Liquid Water Content (LWC), Particle Surface Area (PSA), and the effective droplet radius (re) averaged to 50Hz, 10Hz, and 1Hz.

54 ENVIRONMENTAL SCIENCES↗

Mathematical Morphological Filtering with a Self-Adaptive Reconstruction Technique and Application to Local Seismic Data

Recorded seismic data are generally contaminated by noise from different sources, which masks the signals of interest. In the seismology community, frequency filtering (FF) is the standard method for noise suppression. However, when the signal of interest and noise share the same frequency band, the latter cannot be filtered out without infringing on the former. We implemented a noise suppression approach based on the mathematical morphology theorem. The method involves compound operations of dilation and erosion using structuring elements of varying lengths and decomposes an input noisy waveform into several time functions with differing characteristics. Further, the filtered waveform is constructed from the time functions using a self-adaptive reconstruction technique. Application to a data set of >4700 local waveforms suggests that the implemented mathematical morphological filtering (MMF) approach is efficient for data with low signal-to-noise ratio (SNR) and significantly outperforms FF in that SNR range. For most of the dataset, FF, machine learning (ML) denoising, and continuous wavelet transform (CWT) thresholding result in higher SNR values compared with the MMF method. However, for ~42% of the waveforms, MMF outperforms FF, and the SNR gain achieved with MMF is as large as ~23 dB. Compared to ML denoising and CWT thresholding, this proportion drops to only ~10%–14%. Our results suggests that in an operational setting, MMF cannot replace the other noise suppression methods; however, signal detection can be improved if MMF is used to supplement them in some scenarios. MMF could help detect signals in problematic low-SNR data, which are currently being missed particularly when using FF alone.

58 GEOSCIENCES↗

OPEN-Augmented Reality GUI for Bioenergy Crop Phenotyping and Precision Agriculture (Donald Danforth Plant Science Center Final Scientific Technical Report)

The project led by the Donald Danforth Plant Science Center, in collaboration with Arizona State University, George Washington University, and Saint Louis University, has made significant strides in advancing the phenotypic analysis of bioenergy crops through the development of an innovative AI processing pipeline. This initiative was primarily funded by ARPA-E, with additional cost-sharing provided by the participating institutions. The project successfully utilized a variety of sensors—3D scanners, thermal, RGB, and hyperspectral—to refine algorithms for data-driven trait signature identification and improve the classification and visualization of plant traits. The developed AI processing pipeline is capable of handling the complex, multidimensional data characteristic of dynamic agricultural environments. 1) Contributions to understanding: The research has advanced the field of plant phenomics by showcasing the synergistic use of various sensor data to enhance the precision of trait analysis in bioenergy crops. Through the integration of 3D scanners, thermal, RGB, and hyperspectral sensors, the project has developed robust data-driven trait signature algorithms and visualization techniques. These innovations have facilitated detailed monitoring and management of plant traits, providing vital insights into plant growth dynamics and stress responses. Further, the project has broadened our understanding of how machine learning can be effectively applied in multi-sensor environments to refine trait analysis. By leveraging diverse datasets, the research has not only improved the accuracy of phenotypic assessments but also established a versatile methodological framework that can be extended beyond agriculture to other fields requiring detailed phenotypic analysis. 2) Technical effectiveness and economic feasibility: The AI processing pipeline developed in this project demonstrated significant technical effectiveness, achieving high throughput analysis of extensive phenotypic data and meeting targeted accuracies. This system exemplified the capability of advanced machine learning technologies to efficiently manage and analyze large, complex datasets. Economically, the implementation of the project-developed pipelines may offer substantial cost savings across multiple sectors. It enhances data analysis processes and significantly reduces the need for manual data interpretation, thereby decreasing both the time and resources required. 3) Public benefit: The project has significantly broadened the scope of agricultural methodologies to enhance phenotypic analysis, with potential applications in various sectors beyond agriculture. Additionally, the initiative fostered an enriching educational and collaborative environment, significantly enhancing the technical skills of participants. It also made substantial contributions to the scientific community by providing open-access data sets and tools, encouraging ongoing research and development across various disciplines. Overall, the project not only met its scientific goals but also showcased the extensive utility of integrating advanced machine learning and sensor data analysis technologies. These advancements have proven instrumental in driving forward both theoretical research and practical applications, setting a strong foundation for future explorations and innovations in data-driven science.

60 APPLIED LIFE SCIENCES↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗