Search NASA⌕ Search

SEARCH · Search NASA

Results for “Gap Filling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Technical note: Uncertainties in eddy covariance CO 2 fluxes in a semiarid sagebrush ecosystem caused by gap-filling approaches

Abstract. Gap-filling eddy covariance CO2 fluxes is challenging at dryland sites due to small CO2 fluxes. Here, four machine learning (ML) algorithms including artificial neural network (ANN), k-nearest neighbors (KNNs), random forest (RF), and support vector machine (SVM) are employed and evaluated for gap-filling CO2 fluxes over a semiarid sagebrush ecosystem with different lengths of artificial gaps. The ANN and RF algorithms outperform the KNN and SVM in filling gaps ranging from hours to days, with the RF being more time efficient than the ANN. Performances of the ANN and RF are largely degraded for extremely long gaps of 2 months. In addition, our results suggest that there is no need to fill the daytime and nighttime net ecosystem exchange (NEE) gaps separately when using the ANN and RF. With the ANN and RF, the gap-filling-induced uncertainties in the annual NEE at this site are estimated to be within 16 g C m−2, whereas the uncertainties by the KNN and SVM can be as large as 27 g C m−2. To better fill extremely long gaps of a few months, we test a two-layer gap-filling framework based on the RF. With this framework, the model performance is improved significantly, especially for the nighttime data. Therefore, this approach provides an alternative in filling extremely long gaps to characterize annual carbon budgets and interannual variability in dryland ecosystems.

Yao, Jingyu↗

Gap-filling eddy covariance methane fluxes: Comparison of machine learning model predictions and uncertainties at FLUXNET-CH4 wetlands

Time series of methane fluxes measured by eddy-covariance require gap-filling to estimate annual emissions. Gap-filling methane fluxes is challenging because of high variability and complex responses to multiple drivers. To date, there is no widely established gap-filling standard for methane, with regards both to the best model algorithms and predictors. In this study, we address the need for standardization by synthesizing results of gap-filling methods applied at 17 wetland sites spanning boreal to tropical regions including all major wetlands classes and two rice paddies. We introduce new procedures for: 1) creating realistic artificial gap scenarios, 2) training and evaluating gap-filling models without overstating performance, and 3) predicting half-hourly methane fluxes and annual emissions with robust uncertainty estimates. We tested a conventional method (marginal distribution sampling) and four machine learning algorithms - penalized linear regression, artificial neural networks, random forests, and boosted decision trees - and four predictor sets, including temporal, meteorological, ecosystem carbon and energy flux, and soil predictors. We find that the conventional method can achieve similar median performance to the machine learning models but is worse than the best machine learning models and relatively insensitive to predictor choices. Of the machine learning models, decision tree algorithms performed the best in cross-validation experiments, even with a baseline predictor set, and artificial neural networks showed comparable performance when using all predictors. Soil temperature was frequently the most important predictor whilst water table depth was important at sites with substantial water table fluctuations, highlighting the value of data on soil conditions. Raw gap-filling uncertainties from the machine learning models were underestimated and we propose a method to calibrate uncertainties to observations. Finally, we gap-fill and provide summary evaluation metrics for all 81 sites in the FLUXNET-CH4 community dataset and publicly release the python code for model development, evaluation, and uncertainty estimation.

42 ENGINEERING↗

A novel data gaps filling method for solar PV output forecasting

This study proposes a modified gaps filling method, expanding the column mean imputation method and evaluated using randomly generated missing values comprising 5%, 10%, 15%, and 20% of the original data on power output. The XGBoost algorithm was implemented as a forecasting model using the original and processed datasets and two sources of solar radiation data, namely, Shortwave Radiation (SWR) from Advanced Himawari Imager 8 (AHI-8) and Surface Solar Radiation Downward (SSRD) from ERA5 global reanalysis data. Further, the accuracy of the two sets of forecasted power output was evaluated using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). Results show that by applying the proposed gap filling method and using SWR in forecasting solar photovoltaic (PV) output, the improvement in the RMSE and MAE values range from 12.52% to 24.30% and from 21.10% to 31.31%, respectively. Meanwhile, using SSRD, the improvement in the RMSE values range from 14.01% to 28.54% and MAE values from 22.39% to 35.53%. To further evaluate the accuracy of the proposed gap-filling method, the proposed method could be validated using different datasets and other forecasting methods. Future studies could also consider applying the said method to datasets with data gaps higher than 20%.

Energy & Fuels↗

Detection of surface water temperature variations of Mongolian lakes benefiting from the spatially and temporally gap-filled MODIS data

Lakes provide critical water resources for human activities and ecosystems, particularly in the Mongolian Plateau (MP), which is characterized by a dry climate and a harsh environment. As a region that is sensitive to anthropogenic warming, tracking lake surface water temperature (LSWT) changes in Mongolian lakes is crucial for understanding the consequences of a warming climate on lake ecosystems. However, the long-term monitoring of LSWT is restricted by the spatiotemporal gaps in the raw imagery of remote sensing-based land surface temperature (LST), e.g., the commonly used Moderate Resolution Imaging Spectroradiometer (MODIS) LST products. This study applied an improved gap-filling method by utilizing the discrete cosine transform-based penalized least squares (DCT-PLS) strategy in the spatial domain combined with the linear interpolation (LI) algorithm in the temporal domain. The method was applied to fill gaps in the LSWT imagery of 12 representative lakes across MP. The randomly sampled high-quality MODIS LSWT values in the spatial and temporal domains were excavated as false data gaps and considered “virtual true” validation datasets. The spatial validation results showed that the estimated LSWT for all the lake cases were comparable with the “virtual true” LSWT values, with the average values of the coefficient of determination, mean absolute error, mean square error, and root mean square error being 0.98, 0.38 °C, 0.45 °C, and 0.59 °C, respectively. Meanwhile, the error of nighttime LSWT results was relatively lower than that of daytime LSWT. For temporal interpolation validation, the LI algorithm exhibited relatively better performance and could more objectively indicate the variation in LSWT. Benefiting from the spatially and temporally well-constrained data, we analyzed the interannual and intra-annual change characteristics of the LSWTs of the 12 lakes. The long-term variations of annual and seasonal mean LSWTs in the 12 selected lakes exhibited no evident trends in 2000–2020, while presented apparent interannual fluctuations. The slight changes in the average LSWTs of the 12 selected lakes were in excellent synchronization with the surrounding LST derived from the reanalysis datasets, confirming the widely reported phenomenon of “global warming hiatus” that occurred in the early 21st century. This study improves the understanding of the LSWT variations in Mongolian lakes in response to global climate change. It has the potential to provide an effective approach for monitoring LSWT changes in other large-scale studies.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Assisted Gap-Filled Discharge Data for the East River Community Watershed, Colorado, for Water Years 2014-2021

This dataset contains a collection of machine learning assisted gap-filled discharge data created for all discharge stations across the East River Watershed, Colorado. This data was generated by using raw discharge data collected by Rosemary Carroll, and conducting a random forest machine learning analysis to gap-fill discharge data across all years at the hourly time level. Discharge data with gaps creates problems for analysis of measured and modeled fluxes of carbon and nitrogen exported out of each sub-watershed. Gap-filled data is also required as an input to surface water models, which helps to address our main research question related to how snowmelt timing impacts the timing and magnitude of nitrogen exports. Data is provided in one csv file.

54 ENVIRONMENTAL SCIENCES↗

Gap-filled methane and carbon dioxide fluxes across two ecosystem states at the US-OWC AmeriFlux site (2015−2016, 2020−2022)

This dataset contains gap-filled measurements of methane flux (FCH4), net ecosystem CO2 exchange (NEE) partitioned into gross primary productivity (GPP) and ecosystem respiration (RE), as well as latent heat flux (LE) from a Great Lakes coastal freshwater wetland at the US-OWC AmeriFlux site. The dataset covers the peak growing seasons (June−September) of 2015−2016, dominated by Typha spp., and 2020−2022, characterized by floating-leaved species (lotus and water lily). These data were generated to investigate how rising water levels and vegetation shifts influence CH4 and CO2 fluxes across two distinct ecosystem states in this wetland. The dataset, provided in CSV format, includes half-hourly gap-filled flux data from June to September for 2015, 2016, 2020, 2021, and 2022. The gap-filled data refers to measurements where missing values due to instrument issues or quality control were filled using artificial neural networks (ANNs).

54 ENVIRONMENTAL SCIENCES↗

Hourly gap-filled meteorological data from PIE LTER measurements (2004-2023) used as drivers to run ELM PFLOTRAN simulations

This dataset contains continuous gap-filled precipitation, solar radiation, photosynthetically active radiation (PAR), air temperature, relative humidity, wind speed, and barometric pressure data recorded primarily at the Marshview Farm weather station within the Plum Island Long Term Ecosystems Research (PIE LTER) in Newbury Massachusetts (MA) from 2004 to 2023. We compiled the data set from published annual data packages in 15min resolution available on DataOne. Gaps were filled using different statistical techniques or available observations from the vicinity, e.g. the US-PLo and the US-PHM Ameriflux sites, also located within the PIE LTER. Flags are included in this dataset to indicate the origin of each data point. Metadata files ELMPFLOTRAN_met_dd.csv and ELMPFLOTRAN_met_flmd.csv contain more information on site locations, gap filling protocols, data variables, flags, and QA/QC methods. The data set was used in the spin up and simulations of a land surface model coupled to a biogeochemical reaction network (ELM PFLOTRAN) assessing impacts of hydrology and salinity input on methane fluxes in 2022 and 2023 (Sulman et al., 2024).

54 ENVIRONMENTAL SCIENCES↗

S AP F LOWER : an automated tool for sap flow data preprocessing, gap-filling, and analysis using deep learning

Sap flow, a critical process in plant water use and ecosystem water cycles, is often measured using thermal dissipation probes (TDP) due to their ease of installation and continuous data collection. However, sap flow data frequently include noise, outliers, and gaps, creating challenges for analysis and requiring substantial manual processing. We developed S AP F LOWER , a tool that automates data preprocessing, model training, gap-filling, sapwood area scaling and modeling, and water use analysis. It integrates autocleaning, machine learning and deep learning models (e.g. random forest, Gaussian process regression, long short-term memory (LSTM), bidirectional LSTM (BiLSTM)), and efficient workflows to process sap flow data. S AP F LOWER can remove over 90% of noisy data while preserving legitimate variations and achieve high accuracy in gap-filling based on user-determined parameters. Random forest, LSTM, and BiLSTM models reduced root mean square error to 10% or less for long-term gaps. Model training and prediction can be performed efficiently within seconds. S AP F LOWER significantly enhances the efficiency and accessibility of TDP data analysis by automating complex tasks, enabling researchers without programming expertise to employ advanced techniques. Future improvements will focus on species-specific corrections for TDP and support for additional measurement methods. S AP F LOWER is openly available on GitHub (https://github.com/JiaxinWang123/SapFlower) and Zenodo (doi: 10.5281/zenodo.13665919).

ecosystem water balance↗

15-minute Parker River gap-filled tide height and salinity data, PIE LTER, Plum Island Sound, MA (2014–2023), for ELM PFLOTRAN modeling

This dataset contains 15-minute tide height and salinity data from the Typha site along the Parker River, part of the Plum Island Ecosystems Long Term Ecological Research (PIE LTER) site in Plum Island Sound, Massachusetts (MA) 2014-2023. Tide height (in NAVD88) was compiled from measurements conducted at the mouth of Plum Island Sound and corrected for time lags. Gap-filling of missing periods were done by fitting tidal constituents to the time series. Salinity was measured (and is stored on ESS DIVE ) in 2022 and 2023 using HOBO U24-002 conductivity loggers. River discharge is the most important control on tidal river water salinity at the location (Vallino & Hopkinson, 1998). An artificial neural network was trained to predict river water salinity at the location using Parker River discharge (USGS station 01101000, Parker River at Byfield, MA) and gap-filled salinity observations from a long-term monitoring station ca. 3km downstream from the Typha site (LTER station ‘Middle Road’) as input variables to create continuous time series information. The data set was used in the spin up and simulations of a land surface model coupled to a biogeochemical reaction network (ELM PFLOTRAN) assessing impacts of hydrology and salinity input on methane fluxes in 2022 and 2023 (Sulman et al., 2024). Metadata files ELMPFLOTRAN_tide_salinity_dd.csv and ELMPFLOTRAN_tide_salinity_flmd.csv provide details on site location, data variables, and QA/QC methods .

54 ENVIRONMENTAL SCIENCES↗

A Comparison of Time-Series Gap-Filling Methods to Impute Solar Radiation Data: Preprint

Complete solar resource data sets play a critical role at every stage of solar energy projects; however, measured or modeled solar resource data come with significant uncertainties and usually suffer from several issues, including, but not limited to, data gaps and data quality issues. To mitigate these issues, an appropriate data imputation method should be implemented to build a complete and reliable temporal (and spatial) database. Motivated by this, in this study, we extensively compare the performance of eight different gap-filling methods by creating random and artificial data gaps in (i) hourly irradiance data for 1 year using a few locations of the National Solar Radiation Database (NSRDB) and (ii) 1-minute ground measurement data sets from the Surface Radiation Budget Network (SURFRAD) and the National Renewable Energy Laboratory (NREL) stations.

clearness index↗

A Comparison of Time Series Gap-Filling Methods to Impute Solar Radiation Data

Complete solar resource datasets play a critical role at every stage of solar project phases. However, measured or modeled solar resource data come with significant uncertainties and usually suffer from several issues, including but not limited to, data gaps, data quality issue, etc. In order to mitigate these issues an appropriate data imputation method should be implemented to build a complete and reliable temporal (and spatial) database. Being motivated by this, in this study we compare the performances of eight different gap filling methods extensively by creating random and artificial data gaps in (i) hourly irradiance data for one year using a few locations of the National Solar Radiation Database (NSRDB) and (ii) one-minute ground measurement dataset from Surface Radiation Budget Network (SURFRAD) stations.

clearness index↗

Evaluation of Time-Series Gap-Filling Methods for Solar Irradiance Applications

A complete solar resource data set is essential for any stage of a solar energy project - from feasibility studies to daily operations. But measured or modeled solar resource data are prone to data gaps and data quality issues. To mitigate these issues, a data imputation process should be implemented to obtain a complete and reliable temporal and spatial data series. This study focused on imputing temporal scales by applying random and artificial data gaps and then implementing eight imputation methods, including the Kalman filtering and smoothing and stine interpolations. These methods were implemented on 1-minute to half hourly irradiance data for 1 year using a few locations from the National Solar Radiation Database (NSRDB) and ground measurement data set. The results demonstrated that some of the simpler methods, such as the stine and linear interpolation methods, were the relatively best models based on the statistical metrics for imputing NSRDB and ground measurement data, respectively.

14 SOLAR ENERGY↗

Topologically trivial gap-filling in superconducting Fe(Se,Te) by one-dimensional defects

Abstract Structural distortions and imperfections are a crucial aspect of materials science, on the macroscopic scale providing strength, but also enhancing corrosion and reducing electrical and thermal conductivity. At the nanometre scale, multi-atom imperfections, such as atomic chains and crystalline domain walls have conversely been proposed as a route to topological superconductivity, whose most prominent characteristic is the emergence of Majorana Fermions that can be used for error-free quantum computing. Here, we shed more light on the nature of purported domain walls in Fe(Se,Te) that may host 1D dispersing Majorana modes. We show that the displacement shift of the atomic lattice at these line-defects results from sub-surface impurities that warp the topmost layer(s). Using the electric field between the tip and sample, we manage to reposition the sub-surface impurities, directly visualizing the displacement shift and the underlying defect-free lattice. These results, combined with observations of a completely different type of 1D defect where superconductivity remains fully gapped, highlight the topologically trivial nature of 1D defects in Fe(Se,Te).

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Filling gaps in our understanding of belowground plant traits across the world: an introduction to a Virtual Issue

The belowground world is one of the final frontiers in terrestrial ecology. The tangling of plant roots with the surrounding soil below is a lifeline for the humble forbs and towering trees above, and roots play a key role in shaping ecosystem carbon, water and nutrient cycling (Bardgett et al., 2014). Ecologists have long sought to better understand the ecosystem-scale consequences of differing plant strategies, above- and belowground, by relating plant characteristics, or traits, to plant function (Grime, 1977; Pregitzer, 2002). While developing trait–function linkages is arguably more difficult for plant traits that are hidden belowground, root and rhizosphere ecologists continue to fan out across grasslands and forests with their shovels, isotopes, and specialized cameras, seeking a better understanding of the secret lives of roots. Over the years, New Phytologist has served as a virtual town square for scientists to discuss their hard-won observations on the interplay among belowground plant traits, microbial activity, and edaphic and environmental conditions from biomes around the world (Norby & Jackson, 2000; Pregitzer, 2002; Matamala & Stover, 2013; Norby & Iversen, 2017). In this Editorial we highlight the newest papers that update and add to our understanding of the role of root and rhizosphere traits in broader ecosystem processes. We focused on papers published in New Phytologist between 1 January 2019 and 31 December 2020 and extended this window to papers still in ‘early view’ up to the time of writing. Because of the overwhelming number of papers, we did not include those with a decidedly genomic focus or that served primarily as data syntheses, reviews, insights or commentaries.

59 BASIC BIOLOGICAL SCIENCES↗

Filling gaps in bacterial catabolic pathways with computation and high-throughput genetics

To discover novel catabolic enzymes and transporters, we combined high-throughput genetic data from 29 bacteria with an automated tool to find gaps in their catabolic pathways. GapMind for carbon sources automatically annotates the uptake and catabolism of 62 compounds in bacterial and archaeal genomes. For the compounds that are utilized by the 29 bacteria, we systematically examined the gaps in GapMind’s predicted pathways, and we used the mutant fitness data to find additional genes that were involved in their utilization. We identified novel pathways or enzymes for the utilization of glucosamine, citrulline, myo-inositol, lactose, and phenylacetate, and we annotated 299 diverged enzymes and transporters. We also curated 125 proteins from published reports. For the 29 bacteria with genetic data, GapMind finds high-confidence paths for 85% of utilized carbon sources. In diverse bacteria and archaea, 38% of utilized carbon sources have high-confidence paths, which was improved from 27% by incorporating the fitness-based annotations and our curation. GapMind for carbon sources is available as a web server ( http://papers.genomics.lbl.gov/carbon ) and takes just 30 seconds for the typical genome.

59 BASIC BIOLOGICAL SCIENCES↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Filling the Gaps: A Bayesian Mixture Model for Imputing Missing Soil Water Content Data

ABSTRACT Soil water content (SWC) data are central to evaluating how soil moisture varies over time and space and influences critical plant and ecosystem functions, especially in water‐limited drylands. However, sensors that record SWC at high frequencies often malfunction, leading to incomplete timeseries and limiting our understanding of dryland ecosystem dynamics. We developed an analytical approach to impute missing SWC data, which we tested at six eddy flux tower sites along an elevation gradient in the southwestern United States. We impute missing data as a mixture of linearly interpolated SWC between the observed endpoints of a missing data gap and SWC simulated by an ecosystem water balance model (SOILWAT2). Within a Bayesian framework, we allowed the relative utility (mixture weight) of each component (linearly interpolated vs. SOILWAT2) to vary by depth, site and gap characteristics. We explored “fixed” weights versus “dynamic” weights that vary as a function of cumulative precipitation, average temperature, and time since the start of the gap. Both models estimated missing SWC data well ( R 2 = 0.70–0.88 vs. 0.75–0.91 for fixed vs. dynamic weights, respectively), but the utility of linearly interpolated versus SOILWAT2 values depended on site and depth. SOILWAT2 was more useful for more arid sites, shallower depths, longer and warmer gaps and gaps that received greater precipitation. Overall, the mixture model reliably gap‐fills SWC, while lending insight into processes governing SWC dynamics. This approach to impute missing data could be adapted to accommodate more than two mixture components and other types of environmental timeseries.

Ogle, Kiona [School of Informatics, Computing, and↗