Search NASA⌕ Search

SEARCH · Search NASA

Results for “Missing data problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

On the existence, uniqueness, and asymptotic normality of a consistent solution of the likelihood equations for nonidentically distributed observations: Applications to missing data problems

A general theorem is given which establishes the existence and uniqueness of a consistent solution of the likelihood equations given a sequence of independent random vectors whose distributions are not identical but have the same parameter set. In addition, it is shown that the consistent solution is a MLE and that it is asymptotically normal and efficient. Two applications are discussed: one in which independent observations of a normal random vector have missing components, and the other in which the parameters in a mixture from an exponential family are estimated using independent homogeneous sample blocks of different sizes.

Peters, C.↗

Kriging in the Shadows: Geostatistical Interpolation for Remote Sensing

It is often useful to estimate obscured or missing remotely sensed data. Traditional interpolation methods, such as nearest-neighbor or bilinear resampling, do not take full advantage of the spatial information in the image. An alternative method, a geostatistical technique known as indicator kriging, is described and demonstrated using a Landsat Thematic Mapper image in southern Chiapas, Mexico. The image was first classified into pasture and nonpasture land cover. For each pixel that was obscured by cloud or cloud shadow, the probability that it was pasture was assigned by the algorithm. An exponential omnidirectional variogram model was used to characterize the spatial continuity of the image for use in the kriging algorithm. Assuming a cutoff probability level of 50%, the error was shown to be 17% with no obvious spatial bias but with some tendency to categorize nonpasture as pasture (overestimation). While this is a promising result, the method's practical application in other missing data problems for remotely sensed images will depend on the amount and spatial pattern of the unobscured pixels and missing pixels and the success of the spatial continuity model used.

Rossi, Richard E.↗

Using Concurrent Cardiovascular Information to Augment Survival Time Data from Orthostatic Tilt Tests

Orthostatic Intolerance (OI) is the propensity to develop symptoms of fainting during upright standing. OI is associated with changes in heart rate, blood pressure and other measures of cardiac function. Problem: NASA astronauts have shown increased susceptibility to OI on return from space missions. Current methods for counteracting OI in astronauts include fluid loading and the use of compression garments. Multivariate trajectory spread is greater as OI increases. Pairwise comparisons at the same time within subjects allows incorporation of pass/fail outcomes. Path length, convex hull area, and covariance matrix determinant do well as statistics to summarize this spread Missing data problems Time series analysis need many more time points per OTT session treatment of trend? how incorporate survival information?

Feiveson, Alan H.↗

AUTONOMIE VID

Autonomie Vehicle Information Database (VID) offers a comprehensive list of vehicle specifications since 1990. The database details more than 65,000 vehicles with hundreds of attributes. The database is the result of the development of a general automated data collection framework as well as the development of building blocks for processing, cleaning, integrating and analyzing complex data. The data has undergone several layers of outlier detections processes, machine learning based imputations methods have been used to deal with missing data problems, and new fields have been created according to the rules of feature engineering. Thanks to this streamlined data pipelines, the resulting processed aggregated data should deliver a unique level of information to the user in which the content can be efficiently maintained and updated.

Moswd, Ayman↗

An iterative bidirectional gradient boosting approach for CVR baseline estimation

Here this paper presents a novel Iterative Bidirectional Gradient Boosting Model (IBi-GBM) for estimating the baseline of Conservation Voltage Reduction (CVR) programs. In contrast to many existing methods, we treat CVR baseline estimation as a missing data retrieval problem. The approach involves dividing the load and its corresponding temperature profiles into three periods: pre-CVR, CVR, and post-CVR. To restore the missing load profile during the CVR period, the method employs a three-step process. First, a forward-pass GBM is executed using data from the pre-CVR period as inputs. Subsequently, a backward-pass GBM is applied using data from the post-CVR period. The two restored load profiles are reconciled, considering pre-calculated weights derived from forecasting accuracy, and only the leftmost and rightmost points are retained. The newly restored points are then included as inputs for the subsequent iteration. This iterative procedure continues until the original load data in the CVR period is fully restored. We develop IBi-GBM using actual smart meter and Supervisory Control and Data Acquisition (SCADA) data. Our results demonstrate that IBi-GBM exhibits robust performance across various data resolutions and in different seasons and outperforms existing methods by achieving a 1-2% reduction in normalized Root Mean Square Error (nRMSE).

42 ENGINEERING↗

Spatial Sampling of Weather Data for Regional Crop Yield Simulations

Field-scale crop models are increasingly applied at spatio-temporal scales that range from regions to the globe and from decades up to 100 years. Sufficiently detailed data to capture the prevailing spatio-temporal heterogeneity in weather, soil, and management conditions as needed by crop models are rarely available. Effective sampling may overcome the problem of missing data but has rarely been investigated. In this study the effect of sampling weather data has been evaluated for simulating yields of winter wheat in a region in Germany over a 30-year period (1982-2011) using 12 process-based crop models. A stratified sampling was applied to compare the effect of different sizes of spatially sampled weather data (10, 30, 50, 100, 500, 1000 and full coverage of 34,078 sampling points) on simulated wheat yields. Stratified sampling was further compared with random sampling. Possible interactions between sample size and crop model were evaluated. The results showed differences in simulated yields among crop models but all models reproduced well the pattern of the stratification. Importantly, the regional mean of simulated yields based on full coverage could already be reproduced by a small sample of 10 points. This was also true for reproducing the temporal variability in simulated yields but more sampling points (about 100) were required to accurately reproduce spatial yield variability. The number of sampling points can be smaller when a stratified sampling is applied as compared to a random sampling. However, differences between crop models were observed including some interaction between the effect of sampling on simulated yields and the model used. We concluded that stratified sampling can considerably reduce the number of required simulations. But, differences between crop models must be considered as the choice for a specific model can have larger effects on simulated yields than the sampling strategy. Assessing the impact of sampling soil and crop management data for regional simulations of crop yields is still needed.

upscaling↗

Missing-Data Nonparametric Coherency Estimation

Chave recently proposed an estimator for multitaper spectral density where the time series contains missing values. In this article we generalize this technique to a multitaper estimator of coherence and phase and show that one can also obtain bootstrapped confidence intervals. Additionally, we give two examples. The first is a toy example in which the true coherence is known. In the second example we show that the multitaper missing-data coherence estimator computed on real data with a single gap comprising 11% of the data outperforms the Daniell-smoothed coherence estimator where there are no gaps. The case where the two time series have different missing indices is also discussed.

42 ENGINEERING↗

Reducing a Knowledge-Base Search Space When Data Are Missing

This software addresses the problem of how to efficiently execute a knowledge base in the presence of missing data. Computationally, this is an exponentially expensive operation that without heuristics generates a search space of 1 + 2n possible scenarios, where n is the number of rules in the knowledge base. Even for a knowledge base of the most modest size, say 16 rules, it would produce 65,537 possible scenarios. The purpose of this software is to reduce the complexity of this operation to a more manageable size. The problem that this system solves is to develop an automated approach that can reason in the presence of missing data. This is a meta-reasoning capability that repeatedly calls a diagnostic engine/model to provide prognoses and prognosis tracking. In the big picture, the scenario generator takes as its input the current state of a system, including probabilistic information from Data Forecasting. Using model-based reasoning techniques, it returns an ordered list of fault scenarios that could be generated from the current state, i.e., the plausible future failure modes of the system as it presently stands. The scenario generator models a Potential Fault Scenario (PFS) as a black box, the input of which is a set of states tagged with priorities and the output of which is one or more potential fault scenarios tagged by a confidence factor. The results from the system are used by a model-based diagnostician to predict the future health of the monitored system.

James, Mark↗

Low-rank Tensor Completion for PMU Data Recovery

This paper proposes a tensor completion method for the recovery of missing phasor measurement unit (PMU) measurements. Tensor completion as the general case of matrix completion has attracted increasing attention in recent years. The imputation accuracy for the existing matrix completion methods may be significantly reduced when there are consecutive data losses across multiple data channels. To tackle this issue, we explore the multi-way characteristics of PMU measurements by using a tensor model. We leverage the low-rank property of the PMU measurements and formulate the missing PMU data recovery problem as a low-rank tensor completion problem. An efficient algorithm based on alternating direction method of multipliers (ADMM) is developed to solve the tensor completion problem. The experiments using the real PMU dataset show that the proposed method exhibits better imputation accuracy compared with the conventional data recovery methods.

Ghasemkhani, Amir↗

History and Status of ALSEP and the Apollo Lunar Data Project

A suite of automated scientific instruments (the Apollo Lunar Surface Experiment Package, or ALSEP) was installed at each of the landing sites of Apollo 12, 14, 15, 16, and 17 from 1969 to 1972. They operated from deployment until decommissioning on 30 September 1977. These data were continuously transmitted to Earth and saved on the Range Tapes, which were recorded at the Manned Space Flight Network stations. These data were also broken out by experiment and sent to the experiment Principal Investigators on what were called the P.I. Tapes. Starting in April 1973 the Range Tape data were stored in digital format on 7-track magnetic tapes, the ARCSAV Tapes. In February 1976, the handling of the Range Tapes was transferred to UT Galveston. They produced 9-track tapes referred to as the Work Tapes. Following the Apollo program the Range and ARCSAV tapes, which were never archived, were lost. The Work Tapes were archived at the National Space Science Data Center (NSSDC). Some investigators archived their individual experiment data with NSSDC as well, but much of the data had minimal documentation, were not in digital form, or were stored in difficult to translate formats. Data from many experiments were never delivered to the NSSDC. The Lunar Data Project was started to address the problem of both missing and not readily usable data. Our effort has resulted in recovery of some of the ARCSAV tapes, recovery and digitization of a large volume of Apollo scientific and technical documentation, and restoration of many ALSEP and other Apollo data collections. Restoration involves deciphering formats, assembling necessary ancillary data (metadata), and packaging data in digital format to be archived with the Planetary Data System (PDS). Recovery of the data from the ARCSAV tapes involved having the tapes read on special equipment and extracting the individual experiment data out of the integrated data stream. We will report on the history and status of the various recovery efforts.

Work Tapes↗

State Machine Operation of Complex Systems

Operation of complex systems which depend on one or more other systems with many process variables often operate in more than one state. For each state there may be a variety of parameters of interest, and for each of these, one may require different alarm limits, different archiving needs, and have different critical parameters. Relying on operators to reliably change 10s-1000s of parameters for each system for each state is unreasonable. Not changing these parameters results in alarms being ignored or disabled, critical changes missed, and/or possible data archiving problems. To reliably manage the operation of complex systems, such as cryomodules (CMs), Fermilab is implementing state machines for each CM and an over-arching state machine for the PIP-II superconducting linac (SCL). The state machine transitions and operating parameters are stored/restored to/from a configuration database. Proper implementation of the state machines will not only ensure safe and reliable operation of the CMs, but will help ensure reliable data quality. A description of PIP-II SCL, details of the state machines, and lessons learned from limited use of the state machines in recent CM testing will be discussed.

43 PARTICLE ACCELERATORS↗

Statistical theory and methodology for remote sensing data analysis with special emphasis on LACIE

Crop proportion estimators for determining crop acreage through the use of remote sensing were evaluated. Several studies of these estimators were conducted, including an empirical comparison of the different estimators (using actual data) and an empirical study of the sensitivity (robustness) of the class of mixture estimators. The effect of missing data upon crop classification procedures is discussed in detail including a simulation of the missing data effect. The final problem addressed is that of taking yield data (bushels per acre) gathered at several yield stations and extrapolating these values over some specified large region. Computer programs developed in support of some of these activities are described.

Odell, P. L.↗

Physics-Informed Neural Networks for Heat Transfer Problems

Abstract Physics-informed neural networks (PINNs) have gained popularity across different engineering fields due to their effectiveness in solving realistic problems with noisy data and often partially missing physics. In PINNs, automatic differentiation is leveraged to evaluate differential operators without discretization errors, and a multitask learning problem is defined in order to simultaneously fit observed data while respecting the underlying governing laws of physics. Here, we present applications of PINNs to various prototype heat transfer problems, targeting in particular realistic conditions not readily tackled with traditional computational methods. To this end, we first consider forced and mixed convection with unknown thermal boundary conditions on the heated surfaces and aim to obtain the temperature and velocity fields everywhere in the domain, including the boundaries, given some sparse temperature measurements. We also consider the prototype Stefan problem for two-phase flow, aiming to infer the moving interface, the velocity and temperature fields everywhere as well as the different conductivities of a solid and a liquid phase, given a few temperature measurements inside the domain. Finally, we present some realistic industrial applications related to power electronics to highlight the practicality of PINNs as well as the effective use of neural networks in solving general heat transfer problems of industrial complexity. Taken together, the results presented herein demonstrate that PINNs not only can solve ill-posed problems, which are beyond the reach of traditional computational methods, but they can also bridge the gap between computational and experimental heat transfer.

Engineering↗

QA/QC of the East River, Colorado, discharge and geochemical time series datasets (Almont, BCC, and Pump House) to be used for modeling of hydrogeochemical balance

The following datasets were QA/QC-ed (Quality Assurance/Quality Control): 1. Brush Creek Confluence (BCC) discharge data (from Helen Malenda, USGS, Colorado School of Mines), which were calculated using the pressure transducer data and rating curves. The original 15 min time series data were presented as mean daily discharge. 2. Almont discharge data from United States Geological Survey (USGS). The original data were in 15 min time intervals, and were averaged to mean daily discharge time series. 3. Pump House discharge data as mean daily discharge (downloaded from the SFA portal). 4. BCC and Pump House chemistry data from SFA data portal and/or original spreadsheets provided by Roelof Versteeg. The following challenging QA/QC problems of the datasets were resolved: Missing data with the duration of gaps up to >1 month; Duplicated dates; Anomalies and outliers of discharge and concentrations; Time stamps of measurements of the discharge and concentrations are not aligned (hydrogeochemical balance calculations require the timestamps to be aligned). All QA/QC-ed datasets are given as csv files. The csv files were prepared using the xts files with multiple worksheets, which are also included in the data packages. Figures of the QA/QC-ed datasets are given in the jpeg and pdf formats. The QA/QC-ed datasets have been used to quantify discharge and chemical concentrations in river water in order to understand riverine exports of water and dissolved constituents in the East River watershed. These datasets served as a basis in the presentation given by P. Fox et al. at the 2021 Goldschmidt Conference.

54 ENVIRONMENTAL SCIENCES↗

Restoration of HST images with missing data

Missing data are a fairly common problem when restoring Hubble Space Telescope observations of extended sources. On Wide Field and Planetary Camera images cosmic ray hits and CCD hot spots are the prevalent causes of data losses, whereas on Faint Object Camera images data are lossed due to reseaux marks, blemishes, areas of saturation and the omnipresent frame edges. This contribution discusses a technique for 'filling in' missing data by statistical inference using information from the surrounding pixels. The major gain consists in minimizing adverse spill-over effects to the restoration in areas neighboring those where data are missing. When the mask delineating the support of 'missing data' is made dynamic, cosmic ray hits, etc. can be detected on the fly during restoration.

Adorf, Hans-Martin↗

Case studies in bias reduction and inference for electronic health record data with selection bias and phenotype misclassification

Electronic health records (EHR) are not designed for population‐based research, but they provide easy and quick access to longitudinal health information for a large number of individuals. Many statistical methods have been proposed to account for selection bias, missing data, phenotyping errors, or other problems that arise in EHR data analysis. However, addressing multiple sources of bias simultaneously is challenging. We developed a methodological framework (R package, SAMBA ) for jointly handling both selection bias and phenotype misclassification in the EHR setting that leverages external data sources. These methods assume factors related to selection and misclassification are fully observed, but these factors may be poorly understood and partially observed in practice. As a follow‐up to the methodological work, we demonstrate how to apply these methods for two real‐world case studies, and we evaluate their performance. In both examples, we use individual patient‐level data collected through the University of Michigan Health System and various external population‐based data sources. In case study (a), we explore the impact of these methods on estimated associations between gender and cancer diagnosis. In case study (b), we compare corrected associations between previously identified genetic loci and age‐related macular degeneration with gold standard external summary estimates. These case studies illustrate how to utilize diverse auxiliary information to achieve less biased inference in EHR‐based research.

60 APPLIED LIFE SCIENCES↗

Inpainting radar missing data regions with deep learning

Abstract. Missing and low-quality data regions are a frequent problem for weather radars. They stem from a variety of sources: beam blockage, instrument failure, near-ground blind zones, and many others. Filling in missing data regions is often useful for estimating local atmospheric properties and the application of high-level data processing schemes without the need for preprocessing and error-handling steps – feature detection and tracking, for instance. Interpolation schemes are typically used for this task, though they tend to produce unrealistically spatially smoothed results that are not representative of the atmospheric turbulence and variability that are usually resolved by weather radars. Recently, generative adversarial networks (GANs) have achieved impressive results in the area of photo inpainting. Here, they are demonstrated as a tool for infilling radar missing data regions. These neural networks are capable of extending large-scale cloud and precipitation features that border missing data regions into the regions while hallucinating plausible small-scale variability. In other words, they can inpaint missing data with accurate large-scale features and plausible local small-scale features. This method is demonstrated on a scanning C-band and vertically pointing Ka-band radar that were deployed as part of the Cloud Aerosol and Complex Terrain Interactions (CACTI) field campaign. Three missing data scenarios are explored: infilling low-level blind zones and short outage periods for the Ka-band radar and infilling beam blockage areas for the C-band radar. Two deep-learning-based approaches are tested, a convolutional neural network (CNN) and a GAN that optimize pixel-level error or combined pixel-level error and adversarial loss respectively. Both deep-learning approaches significantly outperform traditional inpainting schemes under several pixel-level and perceptual quality metrics.

54 ENVIRONMENTAL SCIENCES↗