Search NASA⌕ Search

SEARCH · Search NASA

Results for “Missing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Restoration of HST images with missing data

Missing data are a fairly common problem when restoring Hubble Space Telescope observations of extended sources. On Wide Field and Planetary Camera images cosmic ray hits and CCD hot spots are the prevalent causes of data losses, whereas on Faint Object Camera images data are lossed due to reseaux marks, blemishes, areas of saturation and the omnipresent frame edges. This contribution discusses a technique for 'filling in' missing data by statistical inference using information from the surrounding pixels. The major gain consists in minimizing adverse spill-over effects to the restoration in areas neighboring those where data are missing. When the mask delineating the support of 'missing data' is made dynamic, cosmic ray hits, etc. can be detected on the fly during restoration.

Adorf, Hans-Martin↗

Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution

When doing multivariate data analysis, one commonobstacle is the presence of incomplete observations, i.e., observationsfor which one or more key fields are blank. Missing datais often countered by deleting entire observations that containmissing data. The negative effects of deleting entire observationsare multiple: deleting observations reduces sample size andcan also result in biased inferences even if data is missing atrandom. In addition, knowledge contained within incompleteobservations is knowledge lost when they are deleted– and theeffort spent collecting that knowledge is effort wasted. Data imputationmethods, or methods of statistically “filling-in” missingdata, can help combat small sample sizes by using the existinginformation in partially complete observations with the end goalof producing less biased and higher confidence inferences. Whena sample from a multivariate normal population is only partiallycomplete, and the missing data meets appropriate assumptions(missing at random), robust data imputation of the missing datacan be implemented with monotone data augmentation (MDA)using the multivariate t distribution.Missing data imputation is applied to data from the NASA InstrumentCost Model (NICM) using the MDA algorithm underthe assumption of having a multivariate t distribution with fixeddegrees of freedom. A sensitivity analysis to the degrees offreedom parameter is presented to demonstrate robustness ofthe multivariate t distribution when dealing with small samplesas compared to the multivariate normal distribution.

DiNicola, Michael↗

The effects of missing data on global ozone estimates

The effects of missing data and model truncation on estimates of the global mean, zonal distribution, and global distribution of ozone are considered. It is shown that missing data can introduce biased estimates with errors that are not accounted for in the accuracy calculations of empirical modeling techniques. Data-fill techniques are introduced and used for evaluating error bounds and constraining the estimate in areas of sparse and missing data. It is found that the accuracy of the global mean estimate is more dependent on data distribution than model size. Zonal features can be accurately described by 7th order models over regions of adequate data distribution. Data variance accounted for by higher order models appears to represent climatological features of columnar ozone rather than pure error. Data-fill techniques can prevent artificial feature generation in regions of sparse or missing data without degrading high order estimates over dense data regions.

Drewry, J. W.↗

Missing Data and Multiple Imputation: An Unbiased Approach

The default method of dealing with missing data in statistical analyses is to only use the complete observations (complete case analysis), which can lead to unexpected bias when data do not meet the assumption of missing completely at random (MCAR). For the assumption of MCAR to be met, missingness cannot be related to either the observed or unobserved variables. A less stringent assumption, missing at random (MAR), requires that missingness not be associated with the value of the missing variable itself, but can be associated with the other observed variables. When data are truly MAR as opposed to MCAR, the default complete case analysis method can lead to biased results. There are statistical options available to adjust for data that are MAR, including multiple imputation (MI) which is consistent and efficient at estimating effects. Multiple imputation uses informing variables to determine statistical distributions for each piece of missing data. Then multiple datasets are created by randomly drawing on the distributions for each piece of missing data. Since MI is efficient, only a limited number, usually less than 20, of imputed datasets are required to get stable estimates. Each imputed dataset is analyzed using standard statistical techniques, and then results are combined to get overall estimates of effect. A simulation study will be demonstrated to show the results of using the default complete case analysis, and MI in a linear regression of MCAR and MAR simulated data. Further, MI was successfully applied to the association study of CO2 levels and headaches when initial analysis showed there may be an underlying association between missing CO2 levels and reported headaches. Through MI, we were able to show that there is a strong association between average CO2 levels and the risk of headaches. Each unit increase in CO2 (mmHg) resulted in a doubling in the odds of reported headaches.

Foy, M.↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

Optimal Frequency-domain Analysis for Spacecraft Time Series: Introducing the Missing-data Multitaper Power Spectrum Estimator

While the Lomb–Scargle periodogram is foundational to astronomy, it has a significant shortcoming: the variance in the estimated power spectrum does not decrease as more data are acquired. Statisticians have a 60 yr history of developing variance-suppressing power spectrum estimators, but most are not used in astronomy because they are formulated for time series with uniform observing cadence and without seasonal or daily gaps. Here we demonstrate how to apply the missing-data multitaper power spectrum estimator to spacecraft data with uniform time intervals between observations but missing data during thruster fires or momentum dumps. The F-test for harmonic components may be applied to multitaper power spectrum estimates to identify statistically significant oscillations that would not rise above a white noise–based false alarm probability. Multitapering improves the dynamic range of the power spectrum estimate and suppresses spectral window artifacts. We show that the multitaper–F-test combination applied to Kepler observations of KIC 6102338 detects differential rotation without requiring iterative sinusoid fitting and subtraction. Significant signals reside at harmonics of both fundamental rotation frequencies and suggest an antisolar rotation profile. Next we use the missing-data multitaper power spectrum estimator to identify the oscillation modes responsible for the complex "scallop-shell" shape of the K2 light curve of EPIC 203354381. We argue that multitaper power spectrum estimators should be used for all time series with regular observing cadence.

79 ASTRONOMY AND ASTROPHYSICS↗

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

Model certainty in cellular network-driven processes with missing data

Mathematical models are often used to explore network-driven cellular processes from a systems perspective. However, a dearth of quantitative data suitable for model calibration leads to models with parameter unidentifiability and questionable predictive power. Here we introduce a combined Bayesian and Machine Learning Measurement Model approach to explore how quantitative and non-quantitative data constrain models of apoptosis execution within a missing data context. We find model prediction accuracy and certainty strongly depend on rigorous data-driven formulations of the measurement, and the size and make-up of the datasets. For instance, two orders of magnitude more ordinal (e.g., immunoblot) data are necessary to achieve accuracy comparable to quantitative (e.g., fluorescence) data for calibration of an apoptosis execution model. Notably, ordinal and nominal (e.g., cell fate observations) non-quantitative data synergize to reduce model uncertainty and improve accuracy. Finally, we demonstrate the potential of a data-driven Measurement Model approach to identify model features that could lead to informative experimental measurements and improve model predictive power.

59 BASIC BIOLOGICAL SCIENCES↗

Calculation of power spectrums from digital time series with missing data points

Two algorithms are developed for calculating power spectrums from the autocorrelation function when there are missing data points in the time series. Both methods use an average sampling interval to compute lagged products. One method, the correlation function power spectrum, takes the discrete Fourier transform of the lagged products directly to obtain the spectrum, while the other, the modified Blackman-Tukey power spectrum, takes the Fourier transform of the mean lagged products. Both techniques require fewer calculations than other procedures since only 50% to 80% of the maximum lags need be calculated. The algorithms are compared with the Fourier transform power spectrum and two least squares procedures (all for an arbitrary data spacing). Examples are given showing recovery of frequency components from simulated periodic data where portions of the time series are missing and random noise has been added to both the time points and to values of the function. In addition the methods are compared using real data. All procedures performed equally well in detecting periodicities in the data.

Murray, C. W., Jr.↗

Comparing Individualized Survival Predictions From Random Survival Forests and Multistate Models in the Presence of Missing Data: A Case Study of Patients With Oropharyngeal Cancer

Background: In recent years, interest in prognostic calculators for predicting patient health outcomes has grown with the popularity of personalized medicine. These calculators, which can inform treatment decisions, employ many different methods, each of which has advantages and disadvantages. Methods: We present a comparison of a multistate model (MSM) and a random survival forest (RSF) through a case study of prognostic predictions for patients with oropharyngeal squamous cell carcinoma. The MSM is highly structured and takes into account some aspects of the clinical context and knowledge about oropharyngeal cancer, while the RSF can be thought of as a black-box non-parametric approach. Key in this comparison are the high rate of missing values within these data and the different approaches used by the MSM and RSF to handle missingness. Results: We compare the accuracy (discrimination and calibration) of survival probabilities predicted by both approaches and use simulation studies to better understand how predictive accuracy is influenced by the approach to (1) handling missing data and (2) modeling structural/disease progression information present in the data. We conclude that both approaches have similar predictive accuracy, with a slight advantage going to the MSM. Conclusions: Although the MSM shows slightly better predictive ability than the RSF, consideration of other differences are key when selecting the best approach for addressing a specific research question. These key differences include the methods’ ability to incorporate domain knowledge, and their ability to handle missing data as well as their interpretability, and ease of implementation. Ultimately, selecting the statistical method that has the most potential to aid in clinical decisions requires thoughtful consideration of the specific goals.

60 APPLIED LIFE SCIENCES↗

Multitaper Magnitude‐Squared Coherence for Time Series With Missing Data: Understanding Oscillatory Processes Traced by Multiple Observables

To explore the hypothesis of a common source of variability in two time series, observers may estimate the magnitude-squared coherence (MSC), which is a frequency-domain view of the cross correlation. For time series that do not have uniform observing cadence, MSC can be estimated using Welch's overlapping segment averaging. However, multitaper has superior statistical properties to Welch's method in terms of the tradeoff between bias, variance, and bandwidth. The classical multitaper technique has recently been extended to accommodate time series with underlying uniform observing cadence from which some observations are missing. This situation is common for solar and geomagnetic data sets, which may have gaps due to breaks in satellite coverage, instrument downtime, or poor observing conditions. We demonstrate the scientific use of missing-data multitaper magnitude-squared coherence by detecting known solar mid-term oscillations in simultaneous, missing-data time series of solar Lyman α flux and geomagnetic Disturbance Storm Time index. Due to their superior statistical properties, we recommend that multitaper methods be used for all heliospheric time series with underlying uniform observing cadence.

Astro-statistics techniques (1886)↗

Continued Discussion of Failure Mode Modeling and Overall Component Reliability: Are the Data Missing or Censored?

This paper is the continuation of a paper presented at the 13th Probabilistic Safety Assessment and Management Conference, in which a methodology of modeling failure modes of complex components was presented; see Paulos and Smith (2016). This methodology is not particularly helpful in the space industry where there is a lack of failure data, but is more helpful in industries that see a lot of component repairs and improvements, such as in the aircraft or automotive industries. The previous paper demonstrated how the typical approach of treating failure modes as being exponential in nature may yield optimistic predictions when estimating how improvements to components will perform in the future. It is more accurate to model the failure modes as a race in time; unfortunately, this does not give a closed-form solution. This paper uses simulation to solve for the model of the world, and the results compared to the standard methodology of treating the failure modes as being exponential random failures. The standard method is shown to have optimistic predictions, which will lead to prediction errors when failure modes are removed or “fixed.” The failure mode methodology presented in the first paper treated the data as being censored when the test stopped. In this paper, we will compare the results from treating the data as both censored and missing.

Smith, Curtis↗

On the existence, uniqueness, and asymptotic normality of a consistent solution of the likelihood equations for nonidentically distributed observations: Applications to missing data problems

A general theorem is given which establishes the existence and uniqueness of a consistent solution of the likelihood equations given a sequence of independent random vectors whose distributions are not identical but have the same parameter set. In addition, it is shown that the consistent solution is a MLE and that it is asymptotically normal and efficient. Two applications are discussed: one in which independent observations of a normal random vector have missing components, and the other in which the parameters in a mixture from an exponential family are estimated using independent homogeneous sample blocks of different sizes.

Peters, C.↗

Reducing a Knowledge-Base Search Space When Data Are Missing

This software addresses the problem of how to efficiently execute a knowledge base in the presence of missing data. Computationally, this is an exponentially expensive operation that without heuristics generates a search space of 1 + 2n possible scenarios, where n is the number of rules in the knowledge base. Even for a knowledge base of the most modest size, say 16 rules, it would produce 65,537 possible scenarios. The purpose of this software is to reduce the complexity of this operation to a more manageable size. The problem that this system solves is to develop an automated approach that can reason in the presence of missing data. This is a meta-reasoning capability that repeatedly calls a diagnostic engine/model to provide prognoses and prognosis tracking. In the big picture, the scenario generator takes as its input the current state of a system, including probabilistic information from Data Forecasting. Using model-based reasoning techniques, it returns an ordered list of fault scenarios that could be generated from the current state, i.e., the plausible future failure modes of the system as it presently stands. The scenario generator models a Potential Fault Scenario (PFS) as a black box, the input of which is a set of states tagged with priorities and the output of which is one or more potential fault scenarios tagged by a confidence factor. The results from the system are used by a model-based diagnostician to predict the future health of the monitored system.

James, Mark↗