Search NASA⌕ Search

SEARCH · Search NASA

Results for “Missing data problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

On the existence, uniqueness, and asymptotic normality of a consistent solution of the likelihood equations for nonidentically distributed observations: Applications to missing data problems

A general theorem is given which establishes the existence and uniqueness of a consistent solution of the likelihood equations given a sequence of independent random vectors whose distributions are not identical but have the same parameter set. In addition, it is shown that the consistent solution is a MLE and that it is asymptotically normal and efficient. Two applications are discussed: one in which independent observations of a normal random vector have missing components, and the other in which the parameters in a mixture from an exponential family are estimated using independent homogeneous sample blocks of different sizes.

Peters, C.↗

Kriging in the Shadows: Geostatistical Interpolation for Remote Sensing

It is often useful to estimate obscured or missing remotely sensed data. Traditional interpolation methods, such as nearest-neighbor or bilinear resampling, do not take full advantage of the spatial information in the image. An alternative method, a geostatistical technique known as indicator kriging, is described and demonstrated using a Landsat Thematic Mapper image in southern Chiapas, Mexico. The image was first classified into pasture and nonpasture land cover. For each pixel that was obscured by cloud or cloud shadow, the probability that it was pasture was assigned by the algorithm. An exponential omnidirectional variogram model was used to characterize the spatial continuity of the image for use in the kriging algorithm. Assuming a cutoff probability level of 50%, the error was shown to be 17% with no obvious spatial bias but with some tendency to categorize nonpasture as pasture (overestimation). While this is a promising result, the method's practical application in other missing data problems for remotely sensed images will depend on the amount and spatial pattern of the unobscured pixels and missing pixels and the success of the spatial continuity model used.

Rossi, Richard E.↗

Using Concurrent Cardiovascular Information to Augment Survival Time Data from Orthostatic Tilt Tests

Orthostatic Intolerance (OI) is the propensity to develop symptoms of fainting during upright standing. OI is associated with changes in heart rate, blood pressure and other measures of cardiac function. Problem: NASA astronauts have shown increased susceptibility to OI on return from space missions. Current methods for counteracting OI in astronauts include fluid loading and the use of compression garments. Multivariate trajectory spread is greater as OI increases. Pairwise comparisons at the same time within subjects allows incorporation of pass/fail outcomes. Path length, convex hull area, and covariance matrix determinant do well as statistics to summarize this spread Missing data problems Time series analysis need many more time points per OTT session treatment of trend? how incorporate survival information?

Feiveson, Alan H.↗

Spatial Sampling of Weather Data for Regional Crop Yield Simulations

Field-scale crop models are increasingly applied at spatio-temporal scales that range from regions to the globe and from decades up to 100 years. Sufficiently detailed data to capture the prevailing spatio-temporal heterogeneity in weather, soil, and management conditions as needed by crop models are rarely available. Effective sampling may overcome the problem of missing data but has rarely been investigated. In this study the effect of sampling weather data has been evaluated for simulating yields of winter wheat in a region in Germany over a 30-year period (1982-2011) using 12 process-based crop models. A stratified sampling was applied to compare the effect of different sizes of spatially sampled weather data (10, 30, 50, 100, 500, 1000 and full coverage of 34,078 sampling points) on simulated wheat yields. Stratified sampling was further compared with random sampling. Possible interactions between sample size and crop model were evaluated. The results showed differences in simulated yields among crop models but all models reproduced well the pattern of the stratification. Importantly, the regional mean of simulated yields based on full coverage could already be reproduced by a small sample of 10 points. This was also true for reproducing the temporal variability in simulated yields but more sampling points (about 100) were required to accurately reproduce spatial yield variability. The number of sampling points can be smaller when a stratified sampling is applied as compared to a random sampling. However, differences between crop models were observed including some interaction between the effect of sampling on simulated yields and the model used. We concluded that stratified sampling can considerably reduce the number of required simulations. But, differences between crop models must be considered as the choice for a specific model can have larger effects on simulated yields than the sampling strategy. Assessing the impact of sampling soil and crop management data for regional simulations of crop yields is still needed.

upscaling↗

Reducing a Knowledge-Base Search Space When Data Are Missing

This software addresses the problem of how to efficiently execute a knowledge base in the presence of missing data. Computationally, this is an exponentially expensive operation that without heuristics generates a search space of 1 + 2n possible scenarios, where n is the number of rules in the knowledge base. Even for a knowledge base of the most modest size, say 16 rules, it would produce 65,537 possible scenarios. The purpose of this software is to reduce the complexity of this operation to a more manageable size. The problem that this system solves is to develop an automated approach that can reason in the presence of missing data. This is a meta-reasoning capability that repeatedly calls a diagnostic engine/model to provide prognoses and prognosis tracking. In the big picture, the scenario generator takes as its input the current state of a system, including probabilistic information from Data Forecasting. Using model-based reasoning techniques, it returns an ordered list of fault scenarios that could be generated from the current state, i.e., the plausible future failure modes of the system as it presently stands. The scenario generator models a Potential Fault Scenario (PFS) as a black box, the input of which is a set of states tagged with priorities and the output of which is one or more potential fault scenarios tagged by a confidence factor. The results from the system are used by a model-based diagnostician to predict the future health of the monitored system.

James, Mark↗

History and Status of ALSEP and the Apollo Lunar Data Project

A suite of automated scientific instruments (the Apollo Lunar Surface Experiment Package, or ALSEP) was installed at each of the landing sites of Apollo 12, 14, 15, 16, and 17 from 1969 to 1972. They operated from deployment until decommissioning on 30 September 1977. These data were continuously transmitted to Earth and saved on the Range Tapes, which were recorded at the Manned Space Flight Network stations. These data were also broken out by experiment and sent to the experiment Principal Investigators on what were called the P.I. Tapes. Starting in April 1973 the Range Tape data were stored in digital format on 7-track magnetic tapes, the ARCSAV Tapes. In February 1976, the handling of the Range Tapes was transferred to UT Galveston. They produced 9-track tapes referred to as the Work Tapes. Following the Apollo program the Range and ARCSAV tapes, which were never archived, were lost. The Work Tapes were archived at the National Space Science Data Center (NSSDC). Some investigators archived their individual experiment data with NSSDC as well, but much of the data had minimal documentation, were not in digital form, or were stored in difficult to translate formats. Data from many experiments were never delivered to the NSSDC. The Lunar Data Project was started to address the problem of both missing and not readily usable data. Our effort has resulted in recovery of some of the ARCSAV tapes, recovery and digitization of a large volume of Apollo scientific and technical documentation, and restoration of many ALSEP and other Apollo data collections. Restoration involves deciphering formats, assembling necessary ancillary data (metadata), and packaging data in digital format to be archived with the Planetary Data System (PDS). Recovery of the data from the ARCSAV tapes involved having the tapes read on special equipment and extracting the individual experiment data out of the integrated data stream. We will report on the history and status of the various recovery efforts.

Work Tapes↗

Statistical theory and methodology for remote sensing data analysis with special emphasis on LACIE

Crop proportion estimators for determining crop acreage through the use of remote sensing were evaluated. Several studies of these estimators were conducted, including an empirical comparison of the different estimators (using actual data) and an empirical study of the sensitivity (robustness) of the class of mixture estimators. The effect of missing data upon crop classification procedures is discussed in detail including a simulation of the missing data effect. The final problem addressed is that of taking yield data (bushels per acre) gathered at several yield stations and extrapolating these values over some specified large region. Computer programs developed in support of some of these activities are described.

Odell, P. L.↗

Restoration of HST images with missing data

Missing data are a fairly common problem when restoring Hubble Space Telescope observations of extended sources. On Wide Field and Planetary Camera images cosmic ray hits and CCD hot spots are the prevalent causes of data losses, whereas on Faint Object Camera images data are lossed due to reseaux marks, blemishes, areas of saturation and the omnipresent frame edges. This contribution discusses a technique for 'filling in' missing data by statistical inference using information from the surrounding pixels. The major gain consists in minimizing adverse spill-over effects to the restoration in areas neighboring those where data are missing. When the mask delineating the support of 'missing data' is made dynamic, cosmic ray hits, etc. can be detected on the fly during restoration.

Adorf, Hans-Martin↗

Tomographic methods in flow diagnostics

This report presents a viewpoint of tomography that should be well adapted to currently available optical measurement technology as well as the needs of computational and experimental fluid dynamists. The goals in mind are to record data with the fastest optical array sensors; process the data with the fastest parallel processing technology available for small computers; and generate results for both experimental and theoretical data. An in-depth example treats interferometric data as it might be recorded in an aeronautics test facility, but the results are applicable whenever fluid properties are to be measured or applied from projections of those properties. The paper discusses both computed and neural net calibration tomography. The report also contains an overview of key definitions and computational methods, key references, computational problems such as ill-posedness, artifacts, missing data, and some possible and current research topics.

Decker, Arthur J.↗

Holographic interferometry of transparent media with reflection from imbedded test objects

In applying holographic interferometry, opaque objects blocking a portion of the optical beam used to form the interferogram give rise to incomplete data for standard computer tomography algorithms. An experimental technique for circumventing the problem of data blocked by opaque objects is presented. The missing data are completed by forming an interferogram using light backscattered from the opaque object, which is assumed to be diffuse. The problem of fringe localization is considered.

Prikryl, I.↗

An evidential approach to problem solving when a large number of knowledge systems is available

Some recent problems are no longer formulated in terms of imprecise facts, missing data or inadequate measuring devices. Instead, questions pertaining to knowledge and information itself arise and can be phrased independently of any particular area of knowledge. The problem considered in the present work is how to model a problem solver that is trying to find the answer to some query. The problem solver has access to a large number of knowledge systems that specialize in diverse features. In this context, feature means an indicator of what the possibilities for the answer are. The knowledge systems should not be accessed more than once, in order to have truly independent sources of information. Moreover, these systems are allowed to run in parallel. Since access might be expensive, it is necessary to construct a management policy for accessing these knowledge systems. To help in the access policy, some control knowledge systems are available. Control knowledge systems have knowledge about the performance parameters status of the knowledge systems. In order to carry out the double goal of estimating what units to access and to answer the given query, diverse pieces of evidence must be fused. The Dempster-Shafer Theory of Evidence is used to pool the knowledge bases.

Dekorvin, Andre↗

Data Accountability and Uncertainty Analysis for the Mars Science Laboratory

This paper presents machine learning-based approaches to automate and optimize the detection of volume loss for the downlink process of telemetry data from the Mars Curiosity Rover. The Curiosity observes volume loss and data corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the data flow. To resolve this issue, we created a data pipeline to accumulate data from various data sources in the downlink process and detect where the data is missed. In this paper, we benchmarked different methodologies based on the accuracy and excitability of them to identify whether a downlink data that is received to the ground system is complete or incomplete. Our results show that machine learning methods can improve the performance of the GDSA by 55% while the user can diagnose why data is missed and provide an explanation for the data accountability problem.

Chowdhury, Ameera↗

Diagnostic tolerance for missing sensor data

For practical automated diagnostic systems to continue functioning after failure, they must not only be able to diagnose sensor failures but also be able to tolerate the absence of data from the faulty sensors. It is shown that conventional (associational) diagnostic methods will have combinatoric problems when trying to isolate faulty sensors, even if they adequately diagnose other components. Moreover, attempts to extend the operation of diagnostic capability past sensor failure will necessarily compound those difficulties. Model-based reasoning offers a structured alternative that has no special problems diagnosing faulty sensors and can operate gracefully when sensor data is missing.

Scarl, Ethan A.↗

Optimal Codes for the Burst Erasure Channel

Deep space communications over noisy channels lead to certain packets that are not decodable. These packets leave gaps, or bursts of erasures, in the data stream. Burst erasure correcting codes overcome this problem. These are forward erasure correcting codes that allow one to recover the missing gaps of data. Much of the recent work on this topic concentrated on Low-Density Parity-Check (LDPC) codes. These are more complicated to encode and decode than Single Parity Check (SPC) codes or Reed-Solomon (RS) codes, and so far have not been able to achieve the theoretical limit for burst erasure protection. A block interleaved maximum distance separable (MDS) code (e.g., an SPC or RS code) offers near-optimal burst erasure protection, in the sense that no other scheme of equal total transmission length and code rate could improve the guaranteed correctible burst erasure length by more than one symbol. The optimality does not depend on the length of the code, i.e., a short MDS code block interleaved to a given length would perform as well as a longer MDS code interleaved to the same overall length. As a result, this approach offers lower decoding complexity with better burst erasure protection compared to other recent designs for the burst erasure channel (e.g., LDPC codes). A limitation of the design is its lack of robustness to channels that have impairments other than burst erasures (e.g., additive white Gaussian noise), making its application best suited for correcting data erasures in layers above the physical layer. The efficiency of a burst erasure code is the length of its burst erasure correction capability divided by the theoretical upper limit on this length. The inefficiency is one minus the efficiency. The illustration compares the inefficiency of interleaved RS codes to Quasi-Cyclic (QC) LDPC codes, Euclidean Geometry (EG) LDPC codes, extended Irregular Repeat Accumulate (eIRA) codes, array codes, and random LDPC codes previously proposed for burst erasure protection. As can be seen, the simple interleaved RS codes have substantially lower inefficiency over a wide range of transmission lengths.

Hamkins, Jon↗

Reprocessing Microflare Data

The report concerns work on detecting and cataloging solar microflares using an automated. An accompanying figure represents the solar microflare distribution during the period of April 1991 to November 1992, the height of solar activity after the launch of CGRO. It also shows the distribution extending below the distribution obtained at GSFC by manual means. We have implemented significant refinements in the search algorithm. The algorithm in its simplest form searches for transient events and based upon the distribution of the signal among the different BATSE detectors, we can assign it to be of solar origin if the signal distribution conforms to what one expects from a burst or transient from that direction. One of the major problems in an earlier effort was to search for microflares and large flares simultaneously. The requirement for a dynamic range of almost 10(exp 4) resulted in ambiguous identifications at the low side of the distribution. We have since restricted the search to events with peak count rates under 2000/s. Larger events are easily identified in the manual search, so we have chosen not to duplicate that work. The second problem was that missing counts existed below channel 0 in the BATSE Large Area Detector (LAD) data. These have been recovered and are now included in the search process. This provides data below 20 keV, and as we get closer to the thermal part of the spectrum, it provides greater sensitivity. The third problem was that too many BATSE detectors were used in the search. Detectors with pointing directions far from the Sun, although detecting the event, had poorly known responses. Detectors greater than approximately 60 degrees off the Sun are no longer included in the search process. By reducing the systematic errors with the large off-axis detectors we can conduct more rigorous statistical tests of a candidate event to ascertain whether it originated from the solar direction. We have reprocessed the period in the early mission that covers solar maximum and constructed the microflare distribution shown in the figure. The results of the automated search start to deviate from the manual search results below about 1000/s. Not only do we now have this distribution but we have a database of solar microflares that was used to construct the distribution. This database contains the signal at higher energy channels as well as that in channel zero (and below). From this one can, using software at GSFC, construct a photon spectrum for some of the larger microflares. It can also be used in other solar studies, especially those that correlate the X-ray flux with emission at other wavelengths. With some additional effort we hope to integrate this database into the corresponding one residing at the Solar Data Analysis Center at GSFC. The entire CGRO mission's data can now be reprocessed to obtain the microflare distribution at all phases of the solar cycle. This work is in progress. The results of this work will be presented in forthcoming scientific workshops and conferences.

Ryan, James M.↗

Bayesian Analysis of the Cosmic Microwave Background

There is a wealth of cosmological information encoded in the spatial power spectrum of temperature anisotropies of the cosmic microwave background! Experiments designed to map the microwave sky are returning a flood of data (time streams of instrument response as a beam is swept over the sky) at several different frequencies (from 30 to 900 GHz), all with different resolutions and noise properties. The resulting analysis challenge is to estimate, and quantify our uncertainty in, the spatial power spectrum of the cosmic microwave background given the complexities of "missing data", foreground emission, and complicated instrumental noise. Bayesian formulation of this problem allows consistent treatment of many complexities including complicated instrumental noise and foregrounds, and can be numerically implemented with Gibbs sampling. Gibbs sampling has now been validated as an efficient, statistically exact, and practically useful method for low-resolution (as demonstrated on WMAP 1 and 3 year temperature and polarization data). Continuing development for Planck - the goal is to exploit the unique capabilities of Gibbs sampling to directly propagate uncertainties in both foreground and instrument models to total uncertainty in cosmological parameters.

methods - statistical↗

Atmospheric Disturbance Environment Definition

Traditionally, the application of atmospheric disturbance data to airplane design problems has been the domain of the structures engineer. The primary concern in this case is the design of structural components sufficient to handle transient loads induced by the most severe atmospheric "gusts" that might be encountered. The concern has resulted in a considerable body of high altitude gust acceleration data obtained with VGH recorders (airplane velocity, V, vertical acceleration, G, altitude, H) on high-flying airplanes like the U-2 (Ehernberger and Love, 1975). However, the propulsion system designer is less concerned with the accelerations of the airplane than he is with the airflow entering the system's inlet. When the airplane encounters atmospheric turbulence it responds with transient fluctuations in pitch, yaw, and roll angles. These transients, together with fluctuations in the free-stream temperature and pressure will disrupt the total pressure, temperature, Mach number and angularity of the inlet flow. For the mixed compression inlet, the result is a disturbed throat Mach number and/or shock position, and in extreme cases an inlet unstart can occur (cf. Section 2.1). Interest in the effects of inlet unstart on the vehicle dynamics of large, supersonic airplanes is not new. Results published by NASA in 1962 of wind tunnel studies of the problem were used in support of the United States Supersonic Transport program (SST) (White, at aI, 1963). Such studies continued into the late 1970's. However, in spite of such interest, there never was developed an atmospheric disturbance database for inlet unstart analysis to compare with that available for the structures load analysis. Missing were data for the free-stream temperature and pressure disturbances that also contribute to the unStart problem.

Tank, William G.↗