Search NASA⌕ Search

SEARCH · Search NASA

Results for “Missing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Restoration of HST images with missing data

Missing data are a fairly common problem when restoring Hubble Space Telescope observations of extended sources. On Wide Field and Planetary Camera images cosmic ray hits and CCD hot spots are the prevalent causes of data losses, whereas on Faint Object Camera images data are lossed due to reseaux marks, blemishes, areas of saturation and the omnipresent frame edges. This contribution discusses a technique for 'filling in' missing data by statistical inference using information from the surrounding pixels. The major gain consists in minimizing adverse spill-over effects to the restoration in areas neighboring those where data are missing. When the mask delineating the support of 'missing data' is made dynamic, cosmic ray hits, etc. can be detected on the fly during restoration.

Adorf, Hans-Martin↗

Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution

When doing multivariate data analysis, one commonobstacle is the presence of incomplete observations, i.e., observationsfor which one or more key fields are blank. Missing datais often countered by deleting entire observations that containmissing data. The negative effects of deleting entire observationsare multiple: deleting observations reduces sample size andcan also result in biased inferences even if data is missing atrandom. In addition, knowledge contained within incompleteobservations is knowledge lost when they are deleted– and theeffort spent collecting that knowledge is effort wasted. Data imputationmethods, or methods of statistically “filling-in” missingdata, can help combat small sample sizes by using the existinginformation in partially complete observations with the end goalof producing less biased and higher confidence inferences. Whena sample from a multivariate normal population is only partiallycomplete, and the missing data meets appropriate assumptions(missing at random), robust data imputation of the missing datacan be implemented with monotone data augmentation (MDA)using the multivariate t distribution.Missing data imputation is applied to data from the NASA InstrumentCost Model (NICM) using the MDA algorithm underthe assumption of having a multivariate t distribution with fixeddegrees of freedom. A sensitivity analysis to the degrees offreedom parameter is presented to demonstrate robustness ofthe multivariate t distribution when dealing with small samplesas compared to the multivariate normal distribution.

DiNicola, Michael↗

The effects of missing data on global ozone estimates

The effects of missing data and model truncation on estimates of the global mean, zonal distribution, and global distribution of ozone are considered. It is shown that missing data can introduce biased estimates with errors that are not accounted for in the accuracy calculations of empirical modeling techniques. Data-fill techniques are introduced and used for evaluating error bounds and constraining the estimate in areas of sparse and missing data. It is found that the accuracy of the global mean estimate is more dependent on data distribution than model size. Zonal features can be accurately described by 7th order models over regions of adequate data distribution. Data variance accounted for by higher order models appears to represent climatological features of columnar ozone rather than pure error. Data-fill techniques can prevent artificial feature generation in regions of sparse or missing data without degrading high order estimates over dense data regions.

Drewry, J. W.↗

Missing Data and Multiple Imputation: An Unbiased Approach

The default method of dealing with missing data in statistical analyses is to only use the complete observations (complete case analysis), which can lead to unexpected bias when data do not meet the assumption of missing completely at random (MCAR). For the assumption of MCAR to be met, missingness cannot be related to either the observed or unobserved variables. A less stringent assumption, missing at random (MAR), requires that missingness not be associated with the value of the missing variable itself, but can be associated with the other observed variables. When data are truly MAR as opposed to MCAR, the default complete case analysis method can lead to biased results. There are statistical options available to adjust for data that are MAR, including multiple imputation (MI) which is consistent and efficient at estimating effects. Multiple imputation uses informing variables to determine statistical distributions for each piece of missing data. Then multiple datasets are created by randomly drawing on the distributions for each piece of missing data. Since MI is efficient, only a limited number, usually less than 20, of imputed datasets are required to get stable estimates. Each imputed dataset is analyzed using standard statistical techniques, and then results are combined to get overall estimates of effect. A simulation study will be demonstrated to show the results of using the default complete case analysis, and MI in a linear regression of MCAR and MAR simulated data. Further, MI was successfully applied to the association study of CO2 levels and headaches when initial analysis showed there may be an underlying association between missing CO2 levels and reported headaches. Through MI, we were able to show that there is a strong association between average CO2 levels and the risk of headaches. Each unit increase in CO2 (mmHg) resulted in a doubling in the odds of reported headaches.

Foy, M.↗

Calculation of power spectrums from digital time series with missing data points

Two algorithms are developed for calculating power spectrums from the autocorrelation function when there are missing data points in the time series. Both methods use an average sampling interval to compute lagged products. One method, the correlation function power spectrum, takes the discrete Fourier transform of the lagged products directly to obtain the spectrum, while the other, the modified Blackman-Tukey power spectrum, takes the Fourier transform of the mean lagged products. Both techniques require fewer calculations than other procedures since only 50% to 80% of the maximum lags need be calculated. The algorithms are compared with the Fourier transform power spectrum and two least squares procedures (all for an arbitrary data spacing). Examples are given showing recovery of frequency components from simulated periodic data where portions of the time series are missing and random noise has been added to both the time points and to values of the function. In addition the methods are compared using real data. All procedures performed equally well in detecting periodicities in the data.

Murray, C. W., Jr.↗

Continued Discussion of Failure Mode Modeling and Overall Component Reliability: Are the Data Missing or Censored?

This paper is the continuation of a paper presented at the 13th Probabilistic Safety Assessment and Management Conference, in which a methodology of modeling failure modes of complex components was presented; see Paulos and Smith (2016). This methodology is not particularly helpful in the space industry where there is a lack of failure data, but is more helpful in industries that see a lot of component repairs and improvements, such as in the aircraft or automotive industries. The previous paper demonstrated how the typical approach of treating failure modes as being exponential in nature may yield optimistic predictions when estimating how improvements to components will perform in the future. It is more accurate to model the failure modes as a race in time; unfortunately, this does not give a closed-form solution. This paper uses simulation to solve for the model of the world, and the results compared to the standard methodology of treating the failure modes as being exponential random failures. The standard method is shown to have optimistic predictions, which will lead to prediction errors when failure modes are removed or “fixed.” The failure mode methodology presented in the first paper treated the data as being censored when the test stopped. In this paper, we will compare the results from treating the data as both censored and missing.

Smith, Curtis↗

On the existence, uniqueness, and asymptotic normality of a consistent solution of the likelihood equations for nonidentically distributed observations: Applications to missing data problems

A general theorem is given which establishes the existence and uniqueness of a consistent solution of the likelihood equations given a sequence of independent random vectors whose distributions are not identical but have the same parameter set. In addition, it is shown that the consistent solution is a MLE and that it is asymptotically normal and efficient. Two applications are discussed: one in which independent observations of a normal random vector have missing components, and the other in which the parameters in a mixture from an exponential family are estimated using independent homogeneous sample blocks of different sizes.

Peters, C.↗

Reducing a Knowledge-Base Search Space When Data Are Missing

This software addresses the problem of how to efficiently execute a knowledge base in the presence of missing data. Computationally, this is an exponentially expensive operation that without heuristics generates a search space of 1 + 2n possible scenarios, where n is the number of rules in the knowledge base. Even for a knowledge base of the most modest size, say 16 rules, it would produce 65,537 possible scenarios. The purpose of this software is to reduce the complexity of this operation to a more manageable size. The problem that this system solves is to develop an automated approach that can reason in the presence of missing data. This is a meta-reasoning capability that repeatedly calls a diagnostic engine/model to provide prognoses and prognosis tracking. In the big picture, the scenario generator takes as its input the current state of a system, including probabilistic information from Data Forecasting. Using model-based reasoning techniques, it returns an ordered list of fault scenarios that could be generated from the current state, i.e., the plausible future failure modes of the system as it presently stands. The scenario generator models a Potential Fault Scenario (PFS) as a black box, the input of which is a set of states tagged with priorities and the output of which is one or more potential fault scenarios tagged by a confidence factor. The results from the system are used by a model-based diagnostician to predict the future health of the monitored system.

James, Mark↗

Using an Informative Missing Data Model to Predict the Ability to Assess Recovery of Balance Control after Spaceflight

Astronauts show degraded balance control immediately after spaceflight. To assess this change, astronauts' ability to maintain a fixed stance under several challenging stimuli on a movable platform is quantified by "equilibrium" scores (EQs) on a scale of 0 to 100, where 100 represents perfect control (sway angle of 0) and 0 represents data loss where no sway angle is observed because the subject has to be restrained from falling. By comparing post- to pre-flight EQs for actual astronauts vs. controls, we built a classifier for deciding when an astronaut has recovered. Future diagnostic performance depends both on the sampling distribution of the classifier as well as the distribution of its input data. Taking this into consideration, we constructed a predictive ROC by simulation after modeling P(EQ = 0) in terms of a latent EQ-like beta-distributed random variable with random effects.

Feiveson, Alan H.↗

Assessing the Completeness of Occupational Exposure Data in the Lifetime Surveillance of Astronaut Health

INTRODUCTION: Longitudinal analysis on how spaceflight affects human health requires significant amounts of data. Missing data, especially if missing in a non-random fashion, could be a significant challenge to the success and validity of ongoing occupational surveillance and research. Astronaut occupational health data have been collected since 1959 in various formats and as part of several flight programs. As a result of changing methodologies over this span, epidemiologists in the NASA Lifetime Surveillance of Astronaut Health (LSAH) project regularly compile data sets with important exposure or outcome data missing. METHODS: NASA medical records of astronauts participating in voluntary annual LSAH examinations were reviewed and compiled to develop Individual Exposure Profiles (IEP) for each astronaut. These data were supplemented by an interview. If the interview yielded medically relevant information absent from the medical record, that information was considered an update. The IEPs were analyzed to identify trends regarding the characteristics of astronauts who provided updates and what kinds of information were consistently being updated. RESULTS: To date, 190 astronauts have participated in the IEP project. Medical information was updated for 119 individuals during these interviews. The astronauts' likelihood of updating their record upon interview was not significantly related to their spaceflight experience, era of active spaceflight, or duration of longest spaceflight. The most commonly updated categories of medical information were issues encountered during spaceflights, including CO2 symptoms, vision changes, back pain, headaches, and space motion sickness. DISCUSSION: The most commonly updated categories correspond to areas where LSAH has ongoing analysis efforts and therefore do not appear to have been reported at random. This presentation will address identification of missing astronaut health data and trends, forward work identified by the IEP project and how this information may be used for future LSAH data gap analyses.

Sieker, Jeremy↗

Diagnostic tolerance for missing sensor data

For practical automated diagnostic systems to continue functioning after failure, they must not only be able to diagnose sensor failures but also be able to tolerate the absence of data from the faulty sensors. It is shown that conventional (associational) diagnostic methods will have combinatoric problems when trying to isolate faulty sensors, even if they adequately diagnose other components. Moreover, attempts to extend the operation of diagnostic capability past sensor failure will necessarily compound those difficulties. Model-based reasoning offers a structured alternative that has no special problems diagnosing faulty sensors and can operate gracefully when sensor data is missing.

Scarl, Ethan A.↗

NLSI Focus Group on Missing ALSEP Data Recovery: Progress and Plans

On the six Apollo landed missions, the Astronauts deployed the Apollo Lunar Surface Experiments Package (ALSEP) science stations which measured active and passive seismic events, magnetic fields, charged particles, solar wind, heat flow, the diffuse atmosphere, meteorites and their ejecta, lunar dust, etc. Today's scientists are able to extract new information and make new discoveries from the old ALSEP data utilizing recent advances in computer capabilities and new analysis techniques. However, current-day investigators are encountering problems trying to use the ALSEP data. In 2007 archivists from NASA Goddard Space Flight Center (GSFC) National Space Science Data Center (NSSDC) estimated only about 50 percent of the processed ALSEP lunar surface data-of-interest to current lunar science investigators were in the NSSDC archives. The current-day lunar science investigators found most of the ALSEP data, then in the NSSDC archives. were extremely difficult to use. The data were in forms often not well described in the published reports and rerecording anomalies existed in the data which could only be resolved by tape experts. To resolve this problem, the DPS Lunar Data Node was established in 2008 at NSSDC and is in the process of successfully making the existing archived ALSEP data available to current-day investigators in easily useable forms. In July of 2010 the NASA Lunar Science Institute (NLSI) at Ames Research Center established the Recovery of Missing ALSEP Data Focus Group in recognition of the importance of the current activities to find the raw and processed ALSEP data missing from the NSSDC archives.

Lewis, L. R.↗

Managing Large Datasets for Atmospheric Research

Since the mid-1980s, airborne and ground measurements have been widely used to provide comprehensive characterization of atmospheric composition and processes. Field campaigns have generated a wealth of insitu data and have grown considerably over the years in terms of both the number of measured parameters and the data volume. This can largely be attributed to the rapid advances in instrument development and computing power. The users of field data may face a number of challenges spanning data access, understanding, and proper use in scientific analysis. This tutorial is designed to provide an introduction to using data sets, with a focus on airborne measurements, for atmospheric research. The first part of the tutorial provides an overview of airborne measurements and data discovery. This will be followed by a discussion on the understanding of airborne data files. An actual data file will be used to illustrate how data are reported, including the use of data flags to indicate missing data and limits of detection. Retrieving information from the file header will be discussed, which is essential to properly interpreting the data. Field measurements are typically reported as a function of sampling time, but different instruments often have different sampling intervals. To create a combined data set, the data merge process (interpolation of all data to a common time base) will be discussed in terms of the algorithm, data merge products available from airborne studies, and their application in research. Statistical treatment of missing data and data flagged for limit of detection will also be covered in this section. These basic data processing techniques are applicable to both airborne and ground-based observational data sets. Finally, the recently developed Toolsets for Airborne Data (TAD) will be introduced. TAD (tad.larc.nasa.gov) is an airborne data portal offering tools to create user defined merged data products with the capability to provide descriptive statistics and the option to treat measurement uncertainty.

Chen, Gao↗