Search NASA⌕ Search

SEARCH · Search NASA

Results for “Missing Data Recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

Low-rank Tensor Completion for PMU Data Recovery

This paper proposes a tensor completion method for the recovery of missing phasor measurement unit (PMU) measurements. Tensor completion as the general case of matrix completion has attracted increasing attention in recent years. The imputation accuracy for the existing matrix completion methods may be significantly reduced when there are consecutive data losses across multiple data channels. To tackle this issue, we explore the multi-way characteristics of PMU measurements by using a tensor model. We leverage the low-rank property of the PMU measurements and formulate the missing PMU data recovery problem as a low-rank tensor completion problem. An efficient algorithm based on alternating direction method of multipliers (ADMM) is developed to solve the tensor completion problem. The experiments using the real PMU dataset show that the proposed method exhibits better imputation accuracy compared with the conventional data recovery methods.

Ghasemkhani, Amir↗

A Robust Event Diagnostics Platform: Integrating Tensor Analytics and Machine Learning into Real-time Grid Monitoring

The objective of this project is to develop a robust event diagnostics (RED) platform by integrating state-of-the-art tensor analytics and machine learning into real-time grid monitoring. The proposed platform can effectively analyze and discover the information hiding within the provided PMU data for effective real-time grid monitoring. The proposed RED platform provides a set of robust diagnostics tools for grid operation and management, including 1) data quality assessment, 2) data completion, 3) event detection, and 4) robust event classification. All the functionalities of the RED platform can help the operator to make informed decisions and respond in a timely manner. The developed RED platform will serve as an innovative advisory tool to reliably identify key events and discover new insights about the events and grid characteristics in the PMU data, and contribute to the efficient, safe, reliable operation and design of the nation’s electric system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Explaining Missing Data in Graphs: A Constraint-based Approach

Abstract: This paper introduces a constraint-based approach to clarify missing values in graphs. Our method capitalizes on a set S of graph data constraints. An explanation is a sequence of operational enforcement of S towards the recovery of interested yet missing data (e.g., attribute values, edges). We show that constraint-based approach helps us to understand not only why a value is missing, but also how to recover the missing value. We study S-explanation problem, which is to compute the optimal explanations with guarantees on the informativeness and conciseness. We show the problem is in ?P^2 for established graph data constraints such as graph keys and graph association rules. We develop an efficient bidirectional algorithm to compute optimal explanations, without enforcing S on the entire graph. We also show our algorithm can be easily extended to support graph refinement within limited time, and to explain missing answers. Using real-world graphs, we experimentally verify the effectiveness and efficiency of our algorithms.

Data Analytics↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

AlphaFold -assisted structure determination of a bacterial protein of unknown function using X-ray and electron crystallography

Macromolecular crystallography generally requires the recovery of missing phase information from diffraction data to reconstruct an electron-density map of the crystallized molecule. Most recent structures have been solved using molecular replacement as a phasing method, requiring an a priori structure that is closely related to the target protein to serve as a search model; when no such search model exists, molecular replacement is not possible. New advances in computational machine-learning methods, however, have resulted in major advances in protein structure predictions from sequence information. Methods that generate predicted structural models of sufficient accuracy provide a powerful approach to molecular replacement. Taking advantage of these advances, AlphaFold predictions were applied to enable structure determination of a bacterial protein of unknown function (UniProtKB Q63NT7, NCBI locus BPSS0212) based on diffraction data that had evaded phasing attempts using MIR and anomalous scattering methods. Using both X-ray and micro-electron (microED) diffraction data, it was possible to solve the structure of the main fragment of the protein using a predicted model of that domain as a starting point. The use of predicted structural models importantly expands the promise of electron diffraction, where structure determination relies critically on molecular replacement.

molecular replacement↗

Sequential Image Recovery Using Joint Hierarchical Bayesian Learning

Abstract Recovering temporal image sequences (videos) based on indirect, noisy, or incomplete data is an essential yet challenging task. We specifically consider the case where each data set is missing vital information, which prevents the accurate recovery of the individual images. Although some recent (variational) methods have demonstrated high-resolution image recovery based on jointly recovering sequential images, there remain robustness issues due to parameter tuning and restrictions on the type of sequential images. Here, we present a method based on hierarchical Bayesian learning for the joint recovery of sequential images that incorporates prior intra- and inter-image information. Our method restores the missing information in each image by “borrowing” it from the other images. More precisely, we couple sequential images by penalizing their pixel-wise difference. The corresponding penalty terms (one for each pixel and pair of subsequent images) are treated as weakly-informative random variables that favor small pixel-wise differences but allow occasional outliers. As a result, all of the individual reconstructions yield improved accuracy. Our method can be used for various data acquisitions and allows for uncertainty quantification. Some preliminary results indicate its potential use for sequential deblurring and magnetic resonance imaging.

Xiao, Yao↗

Bayesian High-Rank Hankel Matrix Completion for Nonlinear Synchrophasor Data Recovery

Phasor measurement units (PMUs) provide high temporal-resolution synchrophasor measurements for power system monitoring and control. The frequent data quality issues, such as missing and bad data, prevent the incorporation of synchrophasor data in real-time operations. Most existing data-driven data recovery methods assume the power system dynamics can be approximated by a linear dynamical system, and the recovery performance degrades significantly when the power system is experiencing nonlinear dynamics during significant events. Here, this paper proposes a data-driven Bayesian nonlinear synchrophasor data recovery method (Ba-NSDR) that can recover a consecutive time period of simultaneous data losses or errors across all channels, even when the underlying system is highly nonlinear. The idea is to lift the Hankel matrix of the spatial-temporal synchrophasor data to a higher dimension such that the lifted Hankel matrix is low-rank in that space and can be processed with the kernel trick. Our proposed Bayesian method then infers the probabilistic distributions of synchrophasor from the partial observations. Some distinctive features of Ba-NSDR include an uncertainty index to measure the accuracy of the recovery result and the robustness to parameter selections. Our method is verified on both synthetic and recorded event datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A Consensus Equilibrium Approach for 3-D Land Seismic Shots Recovery

Physical and budget constraints often result in inadequate sampling for accurate subsurface imaging. Preprocessing approaches, such as missing trace interpolation, are typically employed to enhance seismic data in such cases. The compressed sensing (CS) framework has been applied for modeling missing seismic data, which is estimated by sparsity-based computational algorithms. While existing work mainly focuses on recovering missing traces resulting from receiver subsampling, source subsampling has greater economical advantages, as sources are more expensive than receivers. Moreover, stronger image models different from sparsity have not been explored for source recovery. This work presents a consensus equilibrium (CE) approach to recover missing seismic shots, which enables to incorporate various regularization operators modeling different data priors. Here, simulation results from a real 3-D land seismic dataset demonstrate that the CE approach provides more accurate estimations of the linear and hyperbolic events in the recovered shots, compared with pure sparsity-based reconstructions.

58 GEOSCIENCES↗

Introduction to Evolve Central Appalachia

The Evolve-CAPP project is developing and implementing strategies to enable the Central Appalachian basin to realize its full economic potential for producing high-value, non-fuel, carbon-ore (CO) products, rare earth elements (REE), and critical minerals (CM). The Basinal Assessment of CORE-CM Resources is a state-of-the-art geological model and comprehensive database that will be used to identify, locate, and quantify in-situ resources in the basin. Restrictive aspects of resource recovery, such as ownership and legal restrictions, are being considered in estimating recoverable resources. A gap analysis of missing or unavailable data that could not be fully integrated in the initial assessment due to budget, access, or other constraints during the assessment effort is planned. A characterization and data acquisition plan to fully characterize the basin’s CORE-CM potential will address gaps in data. The plan will be developed with input from project team members and relevant stakeholders, including land and mineral holding companies and coal operators, who maintain sampling programs as part of their normal operations.

Bishop, Richard↗

Data about data – when, why and how metadata can support the digital plant

A structured approach for recording data quality and contextual information about how and why a signal exists – i.e. metadata – is central to interpret and use sensor data correctly. This is becoming increasingly important with the global trend with data-driven applications such as digital twins and AI-models. But a structured metadata collection and organization of sensor data is not routine in most plants, which can result in lost information and missed opportunities to make use of the investments made in the data collection. Therefore, the IWA task group on Metadata Collection and Organization in wastewater resource recovery systems (MetaCO) was initiated in 2020 and recently delivered the IWA scientific and technical report number 31. The report gives and in-depth description about metadata in water resources recovery facilities (WRRFs) and is available as open access at IWA publishing. The report is the outcome of the collaboration between more than 80 water professionals with the intention to serve WRRF data users with a guide on how to structure and make use of metadata throughout the data pipeline in order to maximize the value of sensor data.

Alferes, Janelcy [VITO, Belgium]↗

Bayesian inference of anisotropic 2D small-angle scattering from sparse measurement

Here, we present a Bayesian inference framework for reconstructing anisotropic two-dimensional small-angle scattering (2D SAS) patterns from sparse, noisy, or partially missing data. The method combines a symmetry-aware angular basis with radial Gaussian process priors to enable accurate, training-free interpolation and denoising. Computational benchmarks demonstrate reliable recovery of both isotropic and high-order anisotropic features under severe data reduction. Experimental validations on stretched polymers, sheared wormlike micelles, and carbon fibers show improved fidelity and resolution compared to raw measurements, achieving comparable accuracy with up to 50-fold fewer detected neutrons. This approach enables quantitative structural analysis under low-flux, time-limited, or single-shot conditions, extending the applicability of 2D SAS techniques to compact neutron sources and mechanically driven soft matter systems undergoing transient structural changes.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Evaluating the Hybrid Modelling Competition: A Step Towards Developing Good Modelling Practice

Hybrid modelling, a combination of mechanistic and data-driven modelling, is a promis¬ing approach to advance current mathematical models towards improved deci¬sion support tools for today's water-related challenges. Researchers have been develop¬ing guidelines or references for good modelling practices in the water field for mecha¬nistic (Rieger et al., 2012) and data-driven (Zhu et al., 2023) modelling, respectively. However, good modelling practices for hybrid modelling are currently missing (Schneider et al., 2022). Therefore, the International Water Association’s (IWA) hybrid modelling working group initiated the first competition on a data science competition platform (i.e. Kaggle) for water resource recovery modelling at the Watermatex con¬ference in September 2023 in Quebec. The main objective of this competition was to gain insights and experience to create good modelling practices. Further goals were to motivate students, researchers, and practitioners model, foster a vibrant and engaged community, and evaluate the efficacy of com¬petitions in solving modelling challenges within the water domain. Our next goal is that facilities will measure and gather relevant data for future competitions to solve their challenges from a modeller’s perspective.

Schneider, Mariane↗

Mixture Model Framework for Traumatic Brain Injury Prognosis Using Heterogeneous Clinical and Outcome Data

Prognoses of Traumatic Brain Injury (TBI) outcomes are neither easily nor accurately determined from clinical indicators. This is due in part to the heterogeneity of damage inflicted to the brain, ultimately resulting in diverse and complex outcomes. Using a data-driven approach on many distinct data elements may be necessary to describe this large set of outcomes and thereby robustly depict the nuanced differences among TBI patients’ recovery. In this work, we develop a method for modeling large heterogeneous data types relevant to TBI. Our approach is geared toward the probabilistic representation of mixed continuous and discrete variables with missing values. The model is trained on a dataset encompassing a variety of data types, including demographics, blood-based biomarkers, and imaging findings. In addition, it includes a set of clinical outcome assessments at 3, 6, and 12 months post-injury. The model is used to stratify patients into distinct groups in an unsupervised learning setting. We use the model to infer outcomes using input data, and show that the collection of input data reduces uncertainty of outcomes over a baseline approach. In addition, we quantify the performance of a likelihood scoring technique that can be used to self-evaluate the extrapolation risk of prognosis on unseen patients.

97 MATHEMATICS AND COMPUTING↗

Deciphering the Role of Total Water Storage Anomalies in Mediating Regional Flooding

Regional floods result from various flood generation mechanisms. Traditional analyses mainly link flooding to extreme rainfall, with limited input from soil moisture. Total water storage (TWS) is a holistic measure of basin wetness, including additional storage components from surface water, snow, and groundwater. Utilizing a new 5-day Gravity Recovery and Climate Experiment and its Follow On (GRACE(-FO)) data set, we investigated the linkage between short-term TWS anomaly (TWSA) and regional flooding. The 5-day TWSA solutions revealed flood signals missed by monthly TWSA solutions. Global basins exhibit distinct storage-discharge co-evolution patterns, offering new insights into flood mechanisms and propensity. Our bivariate event analyses show the annual maximum river discharges co-occur more often with the TWSA maxima than with precipitation in many basins. Further analyses revealed TWSA's time-lagged effect on river discharge, particularly in basins susceptible to floods triggered by saturation-excess runoff. The 5-day TWSA provides a new source of information for enhancing global flood preparedness.

54 ENVIRONMENTAL SCIENCES↗

efam: an e xpanded, metaproteome-supported HMM profile database of viral protein fam ilies

Viruses infect, reprogram and kill microbes, leading to profound ecosystem consequences, from elemental cycling in oceans and soils to microbiome-modulated diseases in plants and animals. Although metagenomic datasets are increasingly available, identifying viruses in them is challenging due to poor representation and annotation of viral sequences in databases. Here, we establish efam, an expanded collection of Hidden Markov Model (HMM) profiles that represent viral protein families conservatively identified from the Global Ocean Virome 2.0 dataset. This resulted in 240 311 HMM profiles, each with at least 2 protein sequences, making efam >7-fold larger than the next largest, pan-ecosystem viral HMM profile database. Adjusting the criteria for viral contig confidence from ‘conservative’ to ‘eXtremely Conservative’ resulted in 37 841 HMM profiles in our efam-XC database. To assess the value of this resource, we integrated efam-XC into VirSorter viral discovery software to discover viruses from less-studied, ecologically distinct oxygen minimum zone (OMZ) marine habitats. This expanded database led to an increase in viruses recovered from every tested OMZ virome by ~24% on average (up to ~42%) and especially improved the recovery of often-missed shorter contigs (<5 kb). Additionally, to help elucidate lesser-known viral protein functions, we annotated the profiles using multiple databases from the DRAM pipeline and virion-associated metaproteomic data, which doubled the number of annotations obtainable by standard, single-database annotation approaches. Together, these marine resources (efam and efam-XC) are provided as searchable, compressed HMM databases that will be updated bi-annually to help maximize viral sequence discovery and study from any ecosystem.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

National Energy Water Treatment & Speciation (NEWTS): A Water & Critical Mineral Database and Dashboard

The scarcity of water resources, the need for beneficial water reuse, and the challenges of wastewater treatment are becoming increasingly pressing in economic, social, and environmental domains. Addressing these concerns requires effective treatment strategies to manage wastewater streams and tackle environmental and economic issues. Furthermore, the recovery of critical minerals from the waste streams associated with energy production holds the promise of offsetting treatment costs and securing local sources of valuable minerals. However, relevant data on these waste streams are dispersed and challenging to locate. The process of ingesting such data into modeling software often involves multiple steps, requiring data restructuring to meet software-input requirements. The non-standardized reporting of water data makes data aggregation and reformatting a time-consuming process. Additionally, essential attributes necessary for modeling water treatment and mineral scale formation are frequently missing. Moreover, data gaps vary depending on the region of interest. Consequently, there is a pressing need for high-quality energy-water composition data that can be easily imported into water chemistry modeling software. To address this need, the National Energy Technology Laboratory has created the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard—a free online tool catering to community leaders and water researchers. NEWTS facilitates a comprehensive understanding of the composition of energy-related wastewater streams in the United States. The datasets provide detailed concentrations and speciation of major and minor aqueous compounds in energy-related wastewater streams, including power plant leachate, acid mine drainage, brackish water, and oil and gas produced water across the United States. Many of the aqueous species are critical minerals (Li, REEs) in high demand to modernize the world’s energy infrastructure. Many of the datasets also contain volumetric flow-rates needed to model the treatment and reuse scenarios in advanced aqueous chemistry software programs. The NEWTS Database and Dashboard offer public access to hitherto challenging-to-access datasets, presented in a standardized format that is tailored for easy input into aqueous chemistry modeling software. By performing the work needed to transform dispersed, disparate data sources into unified, model-ready datasets, NEWTS serves as an essential resource in advancing water treatment research and sustainable water resource management.

produced water management↗

Pennsylvania Department of Environmental Protection (PA DEP) 26r Detailed Produced Water Compositions (version 1.0)

A database of geochemical compositions of aqueous species in produced water reported to the PA DEP. Samples were collected between mid-2012 to early-2020. Data from publicly-available PA DEP 26r reports were scraped from pdf files and cumulated into tabular spreadsheet format for >1000 produced water streams from Marcellus wells in Pennsylvania. In addition to providing the original values, the NETL NEWTS team has reformatted the dataset to allow sample streams to be easily copied into OLI Studio and Geochemist WorkBench (GWB) software for modeling the geochemistry and the recovery of critical minerals, such as lithium, from these produced water streams. In addition, a version of the dataset has been included with predictions for some missing values in the original dataset using machine learning techniques within CoDaRT software, a public ML software developed by the Nation Energy Technology Laboratory. We have made the Input into CoDaRT and one example output from CoDaRT available in this dataset.

Aqueous Chemistry↗