Search NASA⌕ Search

SEARCH · Search NASA

Results for “data imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Short-Term Solar Forecasting Platform Using a Physics-Based Smart Persistence Model and Data Imputation Method

Electrical energy plays vital role in our socio-economic activity and therefore ensuring the reliability of the electric grid, from the generation, transmission and distribution level is critical. In order to maintain the power system parameter viz., frequency, voltage, etc., optimally, balancing of generation and consumption is very much essential. However, solar energy is infirm power by nature this is due to cloud cover / other local phenomena. Hence, Photovoltaic (PV) power generation brings a significant challenge to the grid operator due to the variability of the solar energy. The complexity of this challenge in terms of planning and dispatch ability of PV resources, aggravates with the high penetration of solar energy into the electric grid. In this setting, reliable solar radiation forecasting models based on accurate and quality input data become essential. In order to develop a suitable model for predicting solar radiation, quality historical / real time measurement is also needed. Under this study NIWE and NREL jointly developed / tested short-term solar forecasting frameworks using a smart persistence and physics-based smart persistence models for intra-hour forecasting of solar radiation (PSPI) and benchmarked 9 different data imputation techniques in 15 Solar Radiation Resource Assessment (SRRA) stations, located at different parts of India. During any measurement campaign, due to various technical reasons, we may miss few observations. However, the missing observation often reduce the performance of any forecasting model. Therefore, suitable data imputation method would assist us to obtain continuous observation of solar radiation. A station-by-station and method-by-method analysis was carried out to understand the performance of each model. Based on our analysis, among all the data imputation methods, the Kalman data imputation method is better for Indian Weather condition. In addition, Kalman StructTS, Linear, Stine and Arima methods yield slightly inferior accuracy compared to Kalman, but outperform the other methods. The extended solar radiation data are used by solar forecasting models to provide the prediction of solar radiation at 15 SRRA stations. As far as short term forecasting model is concerned, the PSPI model outperforms the Smart Persistence model. However, the forecast error is increases with the forecasting horizon.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Spatio-Temporal Denoising Graph Autoencoders with Data Augmentation for Missing Photovoltaic Data Imputation

The integration of the global Photovoltaic (PV) market with real time data-loggers has enabled large scale PV data analytical pipelines for power forecasting and long-term reliability assessment of PV fleets. Nevertheless, the performance of PV data analysis heavily depends on the quality of PV timeseries data. This paper proposes a novel Spatio-Temporal Denoising Graph Autoencoder (STD-GAE) framework to impute missing PV Power Data. STDGAE exploits temporal correlation, spatial coherence, and value dependencies from domain knowledge to recover missing data. It is empowered by two modules. (1) To cope with sparse yet various scenarios of missing data, STD-GAE incorporates a domain-knowledge aware data augmentation module that creates plausible variations of missing data patterns. This generalizes STD-GAE to robust imputation over different seasons and environment. (2) STD-GAE nontrivially integrates spatiotemporal graph convolution layers (to recover local missing data by observed “neighboring” PV plants) and denoising autoencoder (to recover corrupted data from augmented counterpart) to improve the accuracy of imputation accuracy at PV fleet level. We have evaluated our proposed model on two realworld PV datasets. Experimental results show that STD-GAE can achieve a gain of 43.14% in imputation accuracy and remains less sensitive to missing rate, different seasons, and missing scenarios, compared with state-of-the-art data imputation methods such as MIDA and LRTC-TNN.

Fan, Yangxin↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗

A Comparison of Time Series Gap-Filling Methods to Impute Solar Radiation Data

Complete solar resource datasets play a critical role at every stage of solar project phases. However, measured or modeled solar resource data come with significant uncertainties and usually suffer from several issues, including but not limited to, data gaps, data quality issue, etc. In order to mitigate these issues an appropriate data imputation method should be implemented to build a complete and reliable temporal (and spatial) database. Being motivated by this, in this study we compare the performances of eight different gap filling methods extensively by creating random and artificial data gaps in (i) hourly irradiance data for one year using a few locations of the National Solar Radiation Database (NSRDB) and (ii) one-minute ground measurement dataset from Surface Radiation Budget Network (SURFRAD) stations.

clearness index↗

A Comparison of Time-Series Gap-Filling Methods to Impute Solar Radiation Data: Preprint

Complete solar resource data sets play a critical role at every stage of solar energy projects; however, measured or modeled solar resource data come with significant uncertainties and usually suffer from several issues, including, but not limited to, data gaps and data quality issues. To mitigate these issues, an appropriate data imputation method should be implemented to build a complete and reliable temporal (and spatial) database. Motivated by this, in this study, we extensively compare the performance of eight different gap-filling methods by creating random and artificial data gaps in (i) hourly irradiance data for 1 year using a few locations of the National Solar Radiation Database (NSRDB) and (ii) 1-minute ground measurement data sets from the Surface Radiation Budget Network (SURFRAD) and the National Renewable Energy Laboratory (NREL) stations.

clearness index↗

A General Spatiotemporal Imputation Framework for Missing Sensor Data

Many applications from precision agriculture, environmental monitoring and transportation networks rely on data collected across space and time over a large geographic area. Missing data poses a significant challenge for any data-driven inference and control tasks. Data imputation or the estimation of missing data can help fill these gaps by utilizing inherent spatial relationships and temporal patterns. A variety of spatiotemporal imputation models have been developed to address missing data in spatiotemporal datasets. However, these classical methods rely on the assumption that the underlying data follows a smooth trend and fail to provide accurate estimates when there is a large number of missing points in the data. Even though there are machine learning driven tensor completion approaches such as convolutional neural network based tensor completion (CoSTCo) that capture the non-linear relationships in the dataset, the transductive nature makes the algorithm less scalable. Thus, existing approaches for estimating the missing information do not effectively capture all dimensions of the spatiotemporal data structure, resulting in erroneous predictions and poor performance. The main contributions of this paper are: (1) We propose a novel inductive framework (G-LSTM) for missing data imputation that integrates a graph neural network with LSTMs to effectively capture both spatial and temporal dependencies. (2) Experimental results on a traffic dataset demonstrate that the proposed GNN integrated with an LSTM framework achieves improved imputation and maintains steady performance even when there are extreme missing conditions in comparison with the state-of-the-art imputation framework (i.e, CoSTCo). (3) The simulation results on a traffic network show up to 69% reduction in mean absolute error and 61% reduction in root mean square error when compared to CoSTCo.

Tharzeen, Aabila↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Evaluation of Time-Series Gap-Filling Methods for Solar Irradiance Applications

A complete solar resource data set is essential for any stage of a solar energy project - from feasibility studies to daily operations. But measured or modeled solar resource data are prone to data gaps and data quality issues. To mitigate these issues, a data imputation process should be implemented to obtain a complete and reliable temporal and spatial data series. This study focused on imputing temporal scales by applying random and artificial data gaps and then implementing eight imputation methods, including the Kalman filtering and smoothing and stine interpolations. These methods were implemented on 1-minute to half hourly irradiance data for 1 year using a few locations from the National Solar Radiation Database (NSRDB) and ground measurement data set. The results demonstrated that some of the simpler methods, such as the stine and linear interpolation methods, were the relatively best models based on the statistical metrics for imputing NSRDB and ground measurement data, respectively.

14 SOLAR ENERGY↗

Detection of Stealthy False Data Injection Attacks in Unobservable Distribution Networks

In this paper, a composite scheme is proposed for detecting stealthy data manipulation attacks on distribution system which is unobservable with standard least squares based state estimators. This technique has three stages where the process of data imputation, voltage phasor estimation and the bad data detection are carried out in a systematic manner. The proposed approach is then integrated with moving target defense strategies which perturbs the network parameters to reveal stealthy false data injection attacks. The proposed approach is tested is validated on a three-phase, unbalanced 37-node distribution system and its results are presented. It is shown that the proposed approach has the ability to accurately detect the presence of FDI attacks using limited measurements (i.e., the test system is unobservable).

Rajasekaran, James K.↗

Generating synthetic occupants for use in building performance simulation

Occupant behaviour simulation frameworks can employ synthetic populations to characterize occupancy and behavioural patterns in buildings based on observed demographic data at a certain geographical location. For buildings, very few synthetic occupant populations have been generated. This paper uses a Bayesian Networks (BN) structural learning approach to synthesize populations of occupants in a multi-family housing case study. Two additional cases of office occupants and senior housing residents are considered as a cross-case comparison. Furthermore, we draw upon the extended version of drivers-needs-actions-systems (DNAS) framework to guide the selection of variables and data imputation. Our results show that the BN approach is powerful in learning the structure of data sets. The synthetic data sets successfully match the joint distributions of the underlying combined data sets. Experiments on the multi-family housing particularly show better performance than the office and senior housing cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine learning models of intermittent operation of RO wellhead water treatment for salinity reduction and nitrate removal

Machine learning models were developed for intermittent multi-mode operation of a wellhead reverse osmosis water purification and desalination system to predict salt passage, nitrate passage, and permeate flux. The models, based on long short-term memory (LSTM) recurrent neural network (RNN) architecture, included an attention mechanism to increase model performance in proximity of the regulatory limit for nitrate. Training and testing of the models for the Startup, Production, Shutdown and Flushing operational modes were based on operational data (consisting of 22 process variables per data sample) acquired every 2–5 s over a six-month period. The significant sets of model input attributes for the different operational modes were assessed via Spearman ranking correlation, Self-Organizing Map (SOM) analysis and feed forward feature selection (FFFS). Although the variability of nitrate passage, salt passage and permeate flux was significant over the four operational modes, prediction performance for the three outcomes were with R2 and Average Absolute Relative Error (AARE) of 0.78–0.95 and 2.96–6.16 %, respectively. Model updates post membrane elements replacement demonstrated similar levels of prediction accuracy. The study results suggest that there is merit in exploring the utility of multi-mode models for sensor fault detection, data imputation, and for potential use in model-predictive control.

Intermittent RO operation↗

Multi Time-scale Imputation aided State Estimation in Distribution System

With the transition to a smart grid, we are witnessing a significant growth in sensor deployments and smart metering infrastructure in the distribution system. However, information from these sensors and meters are typically unevenly sampled at different time-scales and are incomplete. It is critical to effectively aggregate these information sources for situational awareness. In order to reconcile the heterogeneous multi-scale time-series data, we present a multi-task Gaussian process framework. This framework exploits the spatio-temporal correlation across the time-series data to impute data at any desired timescale while providing confidence bounds on the imputations. The value of the imputed data for distribution system operation is illustrated via a matrix completion based state estimation strategy. Results on the IEEE 37 bus distribution system reveals the superior performance of the proposed approach relative to linear interpolation approaches.

Dahale, Shweta↗

TopFusion: Using Topological Feature Space for Fusion and Imputation in Multi-Modal Data

We present a novel multi-modal data fusion technique using topological features. The method, TopFusion, leverages the flexibility of topological data analysis tools (namely persistent homology and persistence images) to map multi-modal datasets into a common feature space by forming a new multi-channel persistence image. Each channel in the image is representative of a view of the data from a modality-dependent filtration. We demonstrate that the topological perspective we take allows for more effective data reconstruction, i.e. imputation. In particular, by performing imputation in topological feature space we are able to outperform the same imputation techniques applied to raw data or alternatively derived features. We show that TopFusion representations can be used as input to downstream deep learning-based computer vision models and doing so achieves comparable performance to other fusion methods for classification on two multi-modal datasets.

Myers, Audun D.↗

Imputation of urban environmental sensor data using gated attention bidirectional long short-term memory (GA-BiLSTM): methods, performance, and implications

Urban environmental monitoring networks frequently encounter significant data gaps due to sensor malfunctions, environmental disturbances, and communication failures. Reliable approaches to address these gaps are essential for ensuring the continuity and quality of environmental data streams. In this study, we developed a gated attention bidirectional long short-term memory (GA-BiLSTM) model to impute missing data in a dense urban monitoring network. Using observations from the CROCUS network in Chicago, we evaluated GA-BiLSTM against widely used approaches (XGBoost and K-nearest neighbors) under scenarios of both short-term intermittent gaps and prolonged outages. GA-BiLSTM consistently outperformed comparative methods, particularly during extended outages of up to ten days, demonstrating its ability to capture spatiotemporal dependencies across sensor nodes. Beyond performance metrics, feature importance and spatial network analyses highlighted the unexpected but critical predictive role of peripheral rural nodes, underlining their strategic value for maintaining robust urban monitoring systems. These results emphasize that advanced imputation methods can substantially improve the reliability of environmental monitoring networks and support more resilient data infrastructures for urban sustainability.

Data imputation↗

Filling the Gaps: A Bayesian Mixture Model for Imputing Missing Soil Water Content Data

ABSTRACT Soil water content (SWC) data are central to evaluating how soil moisture varies over time and space and influences critical plant and ecosystem functions, especially in water‐limited drylands. However, sensors that record SWC at high frequencies often malfunction, leading to incomplete timeseries and limiting our understanding of dryland ecosystem dynamics. We developed an analytical approach to impute missing SWC data, which we tested at six eddy flux tower sites along an elevation gradient in the southwestern United States. We impute missing data as a mixture of linearly interpolated SWC between the observed endpoints of a missing data gap and SWC simulated by an ecosystem water balance model (SOILWAT2). Within a Bayesian framework, we allowed the relative utility (mixture weight) of each component (linearly interpolated vs. SOILWAT2) to vary by depth, site and gap characteristics. We explored “fixed” weights versus “dynamic” weights that vary as a function of cumulative precipitation, average temperature, and time since the start of the gap. Both models estimated missing SWC data well ( R 2 = 0.70–0.88 vs. 0.75–0.91 for fixed vs. dynamic weights, respectively), but the utility of linearly interpolated versus SOILWAT2 values depended on site and depth. SOILWAT2 was more useful for more arid sites, shallower depths, longer and warmer gaps and gaps that received greater precipitation. Overall, the mixture model reliably gap‐fills SWC, while lending insight into processes governing SWC dynamics. This approach to impute missing data could be adapted to accommodate more than two mixture components and other types of environmental timeseries.

Ogle, Kiona [School of Informatics, Computing, and↗

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

High-dimensional data analytics in civil engineering: A review on matrix and tensor decomposition

Recent developments in sensing and monitoring techniques have led to the generation of high-dimensional data in the field of civil engineering. High-dimensional data analytics methods have thus been developed to interpret such complex data. Among the different high-dimensional data analytics techniques, matrix and tensor decomposition methods have acquired a notable interest in the civil engineering community over the past decade. Due to their unique ability to deal with highly redundant and correlated data, these methods are establishing themselves as promising and efficient tools to analyze high-dimensional data in the civil engineering arena. In this paper, high-dimensional data is referred to as a data set in which the number of features is comparable or larger than the number of observations. This review paper aims to summarize the applications of matrix and tensor decomposition methods in civil engineering over the last decade. The survey begins with a general overview of matrix and tensor decomposition followed by highlighting their significance in the field. Afterward, various applications of these high-dimensional data analytics methods in civil engineering are presented, while the advantages offered by these methods are discussed. Lastly, challenges and potential research avenues for employing matrix and tensor decomposition and future emerging trends for their novel use are highlighted.

42 ENGINEERING↗