Search NASA⌕ Search

SEARCH · Search NASA

Results for “Missing data problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

AUTONOMIE VID

Autonomie Vehicle Information Database (VID) offers a comprehensive list of vehicle specifications since 1990. The database details more than 65,000 vehicles with hundreds of attributes. The database is the result of the development of a general automated data collection framework as well as the development of building blocks for processing, cleaning, integrating and analyzing complex data. The data has undergone several layers of outlier detections processes, machine learning based imputations methods have been used to deal with missing data problems, and new fields have been created according to the rules of feature engineering. Thanks to this streamlined data pipelines, the resulting processed aggregated data should deliver a unique level of information to the user in which the content can be efficiently maintained and updated.

Moswd, Ayman↗

An iterative bidirectional gradient boosting approach for CVR baseline estimation

Here this paper presents a novel Iterative Bidirectional Gradient Boosting Model (IBi-GBM) for estimating the baseline of Conservation Voltage Reduction (CVR) programs. In contrast to many existing methods, we treat CVR baseline estimation as a missing data retrieval problem. The approach involves dividing the load and its corresponding temperature profiles into three periods: pre-CVR, CVR, and post-CVR. To restore the missing load profile during the CVR period, the method employs a three-step process. First, a forward-pass GBM is executed using data from the pre-CVR period as inputs. Subsequently, a backward-pass GBM is applied using data from the post-CVR period. The two restored load profiles are reconciled, considering pre-calculated weights derived from forecasting accuracy, and only the leftmost and rightmost points are retained. The newly restored points are then included as inputs for the subsequent iteration. This iterative procedure continues until the original load data in the CVR period is fully restored. We develop IBi-GBM using actual smart meter and Supervisory Control and Data Acquisition (SCADA) data. Our results demonstrate that IBi-GBM exhibits robust performance across various data resolutions and in different seasons and outperforms existing methods by achieving a 1-2% reduction in normalized Root Mean Square Error (nRMSE).

42 ENGINEERING↗

Missing-Data Nonparametric Coherency Estimation

Chave recently proposed an estimator for multitaper spectral density where the time series contains missing values. In this article we generalize this technique to a multitaper estimator of coherence and phase and show that one can also obtain bootstrapped confidence intervals. Additionally, we give two examples. The first is a toy example in which the true coherence is known. In the second example we show that the multitaper missing-data coherence estimator computed on real data with a single gap comprising 11% of the data outperforms the Daniell-smoothed coherence estimator where there are no gaps. The case where the two time series have different missing indices is also discussed.

42 ENGINEERING↗

Low-rank Tensor Completion for PMU Data Recovery

This paper proposes a tensor completion method for the recovery of missing phasor measurement unit (PMU) measurements. Tensor completion as the general case of matrix completion has attracted increasing attention in recent years. The imputation accuracy for the existing matrix completion methods may be significantly reduced when there are consecutive data losses across multiple data channels. To tackle this issue, we explore the multi-way characteristics of PMU measurements by using a tensor model. We leverage the low-rank property of the PMU measurements and formulate the missing PMU data recovery problem as a low-rank tensor completion problem. An efficient algorithm based on alternating direction method of multipliers (ADMM) is developed to solve the tensor completion problem. The experiments using the real PMU dataset show that the proposed method exhibits better imputation accuracy compared with the conventional data recovery methods.

Ghasemkhani, Amir↗

State Machine Operation of Complex Systems

Operation of complex systems which depend on one or more other systems with many process variables often operate in more than one state. For each state there may be a variety of parameters of interest, and for each of these, one may require different alarm limits, different archiving needs, and have different critical parameters. Relying on operators to reliably change 10s-1000s of parameters for each system for each state is unreasonable. Not changing these parameters results in alarms being ignored or disabled, critical changes missed, and/or possible data archiving problems. To reliably manage the operation of complex systems, such as cryomodules (CMs), Fermilab is implementing state machines for each CM and an over-arching state machine for the PIP-II superconducting linac (SCL). The state machine transitions and operating parameters are stored/restored to/from a configuration database. Proper implementation of the state machines will not only ensure safe and reliable operation of the CMs, but will help ensure reliable data quality. A description of PIP-II SCL, details of the state machines, and lessons learned from limited use of the state machines in recent CM testing will be discussed.

43 PARTICLE ACCELERATORS↗

Physics-Informed Neural Networks for Heat Transfer Problems

Abstract Physics-informed neural networks (PINNs) have gained popularity across different engineering fields due to their effectiveness in solving realistic problems with noisy data and often partially missing physics. In PINNs, automatic differentiation is leveraged to evaluate differential operators without discretization errors, and a multitask learning problem is defined in order to simultaneously fit observed data while respecting the underlying governing laws of physics. Here, we present applications of PINNs to various prototype heat transfer problems, targeting in particular realistic conditions not readily tackled with traditional computational methods. To this end, we first consider forced and mixed convection with unknown thermal boundary conditions on the heated surfaces and aim to obtain the temperature and velocity fields everywhere in the domain, including the boundaries, given some sparse temperature measurements. We also consider the prototype Stefan problem for two-phase flow, aiming to infer the moving interface, the velocity and temperature fields everywhere as well as the different conductivities of a solid and a liquid phase, given a few temperature measurements inside the domain. Finally, we present some realistic industrial applications related to power electronics to highlight the practicality of PINNs as well as the effective use of neural networks in solving general heat transfer problems of industrial complexity. Taken together, the results presented herein demonstrate that PINNs not only can solve ill-posed problems, which are beyond the reach of traditional computational methods, but they can also bridge the gap between computational and experimental heat transfer.

Engineering↗

The Missing Satellite Problem outside of the Local Group. II. Statistical Properties of Satellites of Milky Way–like Galaxies

We present a new observation of satellite galaxies around seven Milky Way (MW)–like galaxies located outside of the Local Group (LG) using Subaru/Hyper Suprime-Cam imaging data to statistically address the missing satellite problem. We select satellite galaxy candidates using magnitude, surface brightness, Sérsic index, axial ratio, FWHM, and surface brightness fluctuation cuts, followed by visual screening of false positives such as optical ghosts of bright stars. We identify 51 secure dwarf satellite galaxies within the virial radius of nine host galaxies, two of which are drawn from the pilot observation presented in Paper I. We find that the average luminosity function of the satellite galaxies is consistent with that of the MW satellites, although the luminosity function of each host galaxy varies significantly. We observe an indication that more massive hosts tend to have a larger number of satellites. Physical properties of the satellites such as the size–luminosity relation are also consistent with the MW satellites. However, the spatial distribution is different; we find that the satellite galaxies outside of the LG show no sign of concentration or alignment, while that of the MW satellites is more concentrated around the host and exhibits a significant alignment. As we focus on relatively massive satellites with M V < –10, we do not expect that the observational incompleteness can be responsible here. This trend might represent a peculiarity of the MW satellites, and further work is needed to understand its origin.

79 ASTRONOMY AND ASTROPHYSICS↗

QA/QC of the East River, Colorado, discharge and geochemical time series datasets (Almont, BCC, and Pump House) to be used for modeling of hydrogeochemical balance

The following datasets were QA/QC-ed (Quality Assurance/Quality Control): 1. Brush Creek Confluence (BCC) discharge data (from Helen Malenda, USGS, Colorado School of Mines), which were calculated using the pressure transducer data and rating curves. The original 15 min time series data were presented as mean daily discharge. 2. Almont discharge data from United States Geological Survey (USGS). The original data were in 15 min time intervals, and were averaged to mean daily discharge time series. 3. Pump House discharge data as mean daily discharge (downloaded from the SFA portal). 4. BCC and Pump House chemistry data from SFA data portal and/or original spreadsheets provided by Roelof Versteeg. The following challenging QA/QC problems of the datasets were resolved: Missing data with the duration of gaps up to >1 month; Duplicated dates; Anomalies and outliers of discharge and concentrations; Time stamps of measurements of the discharge and concentrations are not aligned (hydrogeochemical balance calculations require the timestamps to be aligned). All QA/QC-ed datasets are given as csv files. The csv files were prepared using the xts files with multiple worksheets, which are also included in the data packages. Figures of the QA/QC-ed datasets are given in the jpeg and pdf formats. The QA/QC-ed datasets have been used to quantify discharge and chemical concentrations in river water in order to understand riverine exports of water and dissolved constituents in the East River watershed. These datasets served as a basis in the presentation given by P. Fox et al. at the 2021 Goldschmidt Conference.

54 ENVIRONMENTAL SCIENCES↗

Case studies in bias reduction and inference for electronic health record data with selection bias and phenotype misclassification

Electronic health records (EHR) are not designed for population‐based research, but they provide easy and quick access to longitudinal health information for a large number of individuals. Many statistical methods have been proposed to account for selection bias, missing data, phenotyping errors, or other problems that arise in EHR data analysis. However, addressing multiple sources of bias simultaneously is challenging. We developed a methodological framework (R package, SAMBA ) for jointly handling both selection bias and phenotype misclassification in the EHR setting that leverages external data sources. These methods assume factors related to selection and misclassification are fully observed, but these factors may be poorly understood and partially observed in practice. As a follow‐up to the methodological work, we demonstrate how to apply these methods for two real‐world case studies, and we evaluate their performance. In both examples, we use individual patient‐level data collected through the University of Michigan Health System and various external population‐based data sources. In case study (a), we explore the impact of these methods on estimated associations between gender and cancer diagnosis. In case study (b), we compare corrected associations between previously identified genetic loci and age‐related macular degeneration with gold standard external summary estimates. These case studies illustrate how to utilize diverse auxiliary information to achieve less biased inference in EHR‐based research.

60 APPLIED LIFE SCIENCES↗

Inpainting radar missing data regions with deep learning

Abstract. Missing and low-quality data regions are a frequent problem for weather radars. They stem from a variety of sources: beam blockage, instrument failure, near-ground blind zones, and many others. Filling in missing data regions is often useful for estimating local atmospheric properties and the application of high-level data processing schemes without the need for preprocessing and error-handling steps – feature detection and tracking, for instance. Interpolation schemes are typically used for this task, though they tend to produce unrealistically spatially smoothed results that are not representative of the atmospheric turbulence and variability that are usually resolved by weather radars. Recently, generative adversarial networks (GANs) have achieved impressive results in the area of photo inpainting. Here, they are demonstrated as a tool for infilling radar missing data regions. These neural networks are capable of extending large-scale cloud and precipitation features that border missing data regions into the regions while hallucinating plausible small-scale variability. In other words, they can inpaint missing data with accurate large-scale features and plausible local small-scale features. This method is demonstrated on a scanning C-band and vertically pointing Ka-band radar that were deployed as part of the Cloud Aerosol and Complex Terrain Interactions (CACTI) field campaign. Three missing data scenarios are explored: infilling low-level blind zones and short outage periods for the Ka-band radar and infilling beam blockage areas for the C-band radar. Two deep-learning-based approaches are tested, a convolutional neural network (CNN) and a GAN that optimize pixel-level error or combined pixel-level error and adversarial loss respectively. Both deep-learning approaches significantly outperform traditional inpainting schemes under several pixel-level and perceptual quality metrics.

54 ENVIRONMENTAL SCIENCES↗

Effective Missing Value Imputation Methods for Building Monitoring Data

To understand behaviors of natural and man-made events, such as energy consumption of buildings, which accounts for 40% of energy uses in the US, we deploy automated monitoring devices to record periodic observations. However, such experimental and observation data often contains problems and irregularities that have to be cleaned up before analyses. Due to various conditions affecting sensor operations, the communication channels, recording steps, or the recording media, the recorded data might have missing values, errors, or anomalous values. An effective way to clean up these problems is to replace these missing values, errors and anomalous values with expected values, a process generally known as imputation. In this work, we survey commonly used missing value imputation techniques and compare their performance on a set of building monitoring data. To compare the different types of sensor measurements with widely varying characteristics, we use normalized root mean squared error (NRMSE) as the key metric for the effectiveness of the imputation methods. We additionally consider periodicity and run time when considering comparing methods. Through extensive testing, we find that for small gap sizes, up to 8 consecutive missing values, linear interpolation performs the best; for larger gaps stretching up to 48 consecutive missing values, K-nearest neighbors provides the most accurate imputations; for even larger gaps, more computational intensive methods, such as matrix factorization, achieve the smallest NRMSE. Additionally, we observe that these computationally intensive algorithms not only provide accurate imputations for large gaps, but are also more robust across all types of sensors.

Cho, B↗

Identifying recharge sources and their impacts on a North Central New Mexico shallow aquifer using unsupervised machine learning

In this article, shallow aquifers are important but highly variable resources in arid to semi-arid regions. Limited shallow aquifer volume results in high sensitivity to recharge fluctuations, which can impact the local fauna and flora, and transport of contaminants in the aquifer or vadose zone. Aquifer response to external forcing (e.g., precipitation) is usually solved by estimating aquifer parameters and running physics-based models to match known fluctuations of hydraulic head. However, this technique is time and computationally expensive. Furthermore, high aquifer complexity decreases precision in physics-based models. Alternatively supervised machine learning is used to predict aquifer dynamics. However, these techniques rely on input data and struggle to interpret aquifer response for missing sources (i.e., snowpack data). To counter these problems, we propose an unsupervised machine learning technique (NMFk) to estimate the impact of different sources on aquifer recharge. NMFk is used to understand the influence of external forcing on shallow aquifer recharge in the Pajarito Plateau (Los Alamos, NM, USA). The results show how NMFk can be used to reduce the data dimension in a complex field dataset to three recharge signals that cause fluctuations within the field data. Here, the source signals are interpreted as rainfall, snowmelt, and a delayed aquifer response to the previous two signals. These results evidence how heterogeneous aquifers delimited by canyons incised into the Pajarito Plateau respond in similar ways to the source signals identified by NMFk. Furthermore, results show the importance of the local geology where faults act as sinks, and anthropogenic disturbances can facilitate infiltration amplifying the interpreted signal.

54 ENVIRONMENTAL SCIENCES↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

AGR 5/6/7 Data Qualification Report for ATR Cycles 162B through 168A

This report provides the qualification status of experimental data for the Advanced Gas Reactor (AGR) 5/6/7 fuel irradiation. AGR-5/6/7 was conducted in the Advanced Test Reactor (ATR) at Idaho National Laboratory (INL) in support of development and qualification of tri-structural isotropic (TRISO) low-enriched fuel for use in high temperature gas-cooled reactors. The objectives of the AGR-5/6/7 experiments are to: (i) irradiate reference-design fuel particles to support fuel qualification, (ii) establish operating margins for the fuel beyond normal operating conditions, and (iii) provide irradiated-fuel performance data and irradiated-fuel samples for post-irradiation examination (PIE) and safety testing. The test train contains five separate capsules that were independently controlled and monitored. Each capsule contains multiple 12.51-mm-long compacts filled with low enriched uranium carbide/oxide (UCO) TRISO fuel particles. The primary objective of the AGR-5/6 test (Capsules 1, 2, 4, and 5) is to verify successful performance of the reference-design fuel under normal operating conditions. The AGR-7 test (Capsule 3) was designed to explore fuel performance at higher temperatures to demonstrate the capability of the fuel to withstand conditions beyond normal operating conditions in support of plant design and licensing. AGR 5/6/7 will also provide irradiated-fuel performance data on fission-gas release from failed particles during irradiation. The AGR-5/6/7 capsules were irradiated in the ATR northeast flux trap location. The experiment began on February 16, 2018 and ended on July 22, 2020, spanning nine ATR cycles over two and a half years. Thus, the AGR-5/6/7 fuel compacts were irradiated for a total of 360.9 effective full power days. The AGR 5/6/7 experiment was able to remain in the reactor core during all three Powered Axial Locator Mechanism (PALM) cycles (163A, 165A, and 167A) without overheating its fuel compacts. This report includes irradiation monitoring data from nine ATR Cycles: 162B, 163A, 164A, 164B, 165A, 166A, 166B, 167A, and 168A, as stored in the Nuclear Data Management and Analysis System (NDMAS). During irradiation, data records consisted of instantaneous measurements recorded every minute and provided by text files automatically every 2 hours. The AGR 5/6/7 data streams addressed in this report include thermocouple (TC) temperatures, sweep gas data (flow rates [capsule inlet, outlet, and downstream at detector], pressure, and moisture content), and Fission Product Monitoring System (FPMS) data (release rates and release to birth rate ratios [R/Bs]) for each of the five capsules. A total of 94,989,908 TC temperature and sweep gas data records were received and processed by NDMAS for AGR 5/6/7 irradiation. Of these records, 41,593,387 (or 43.7% of the total) met data collection and accuracy requirements and are labeled as Qualified. A total of 57,746,693 TC temperature readings were captured from 54 installed TCs. Among them, 10,034,676 TC temperature records (only 17.4%) were Qualified and 47,701,371 TC temperatures (or 82.6%) are Failed due to 48 TC failures (63.5%) and due to missing values (19.1%). To assess performance of the operational TCs, analysis of daily correlations between TCs found no evidence of virtual junction failure for any TCs. Analyses on control charts of TC temperature differences revealed trending in TC readings for TC2, 4, 5, and 13 in Capsule 3, but there is no conclusive indication of TC drift failure that caused those trends. Therefore, TC control charts are not used to disqualify TC data, but only for users’ consideration. For sweep gas flow rates, a total of 31,519,747 gas flow records (84.4%) are Qualified for use for AGR-5/6/7 experiment; 5,723,468 gas flow records (15.4%) are Failed due mostly to missing values; and 74,641 high sweep gas flow rates (0.2 %) are Trend. A large number of Failed missing TC temperature and gas flow values were caused by an error in the data output script that outputted a ‘NULL’ value when values were unchanged. This problem was fixed during the outage of Cycle 166B, which led to a substantially decreased number of missing values during the last three cycles. Nonetheless, a large amount of non-missing data remained because of the high data acquisition frequency (1-minute) and still provided sufficient data to effectively monitor the experiment as designed. For FPMS data, NDMAS received and processed fission product release and R/B data for nine ATR cycles, when ATR core reached full power during AGR 5/6/7 irradiation. These data consist of 110,388 release rate records and 110,388 R/B records for the twelve radionuclides (Kr 85m, Kr 87, Kr 88, Kr 89, Kr 90, Xe 131m, Xe 133, Xe 135, Xe 135m, Xe 137, Xe 138, and Xe 139) for each of the five capsules. Equivalent numbers of uncertainty records associated the release rates and R/B values were provided. To date, qualification status of the FPMS data stored in the NDMAS dat

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗

Statistical Performance of Forced Oscillation Detectors in the Presence of Missing Measurements

In bulk power systems, measurement-based monitoring for large oscillations can help maintain system reliability. One of the challenges encountered in a recent field demonstration was the unavailability of measurements due to underlying measurement quality or communication problems. During the demonstration, the oscillation detector ignored a measurement location if even 10 seconds of data was missing. To extend the detector's ability to operate in these conditions, this paper evaluates the impact of three methods for addressing missing data. The strengths and weaknesses of each approach are evaluated using theoretical expressions for the probability of detection along with results from simulated data and publicly available field measurements. Based on these results, a suitable approach is identified that can extend the oscillation detector's performance when large segments of data are missing.

Follum, James D.↗

Explaining Missing Data in Graphs: A Constraint-based Approach

Abstract: This paper introduces a constraint-based approach to clarify missing values in graphs. Our method capitalizes on a set S of graph data constraints. An explanation is a sequence of operational enforcement of S towards the recovery of interested yet missing data (e.g., attribute values, edges). We show that constraint-based approach helps us to understand not only why a value is missing, but also how to recover the missing value. We study S-explanation problem, which is to compute the optimal explanations with guarantees on the informativeness and conciseness. We show the problem is in ?P^2 for established graph data constraints such as graph keys and graph association rules. We develop an efficient bidirectional algorithm to compute optimal explanations, without enforcing S on the entire graph. We also show our algorithm can be easily extended to support graph refinement within limited time, and to explain missing answers. Using real-world graphs, we experimentally verify the effectiveness and efficiency of our algorithms.

Data Analytics↗