Search NASA⌕ Search

SEARCH · Search NASA

Results for “Missing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Missing data in multi-omics integration: Recent advances through artificial intelligence

Biological systems function through complex interactions between various ‘omics (biomolecules), and a more complete understanding of these systems is only possible through an integrated, multi-omic perspective. This has presented the need for the development of integration approaches that are able to capture the complex, often non-linear, interactions that define these biological systems and are adapted to the challenges of combining the heterogenous data across ‘omic views. A principal challenge to multi-omic integration is missing data because all biomolecules are not measured in all samples. Due to either cost, instrument sensitivity, or other experimental factors, data for a biological sample may be missing for one or more ‘omic techologies. Recent methodological developments in artificial intelligence and statistical learning have greatly facilitated the analyses of multi-omics data, however many of these techniques assume access to completely observed data. A subset of these methods incorporate mechanisms for handling partially observed samples, and these methods are the focus of this review. We describe recently developed approaches, noting their primary use cases and highlighting each method's approach to handling missing data. We additionally provide an overview of the more traditional missing data workflows and their limitations; and we discuss potential avenues for further developments as well as how the missing data issue and its current solutions may generalize beyond the multi-omics context.

97 MATHEMATICS AND COMPUTING↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

Optimal Frequency-domain Analysis for Spacecraft Time Series: Introducing the Missing-data Multitaper Power Spectrum Estimator

While the Lomb–Scargle periodogram is foundational to astronomy, it has a significant shortcoming: the variance in the estimated power spectrum does not decrease as more data are acquired. Statisticians have a 60 yr history of developing variance-suppressing power spectrum estimators, but most are not used in astronomy because they are formulated for time series with uniform observing cadence and without seasonal or daily gaps. Here we demonstrate how to apply the missing-data multitaper power spectrum estimator to spacecraft data with uniform time intervals between observations but missing data during thruster fires or momentum dumps. The F-test for harmonic components may be applied to multitaper power spectrum estimates to identify statistically significant oscillations that would not rise above a white noise–based false alarm probability. Multitapering improves the dynamic range of the power spectrum estimate and suppresses spectral window artifacts. We show that the multitaper–F-test combination applied to Kepler observations of KIC 6102338 detects differential rotation without requiring iterative sinusoid fitting and subtraction. Significant signals reside at harmonics of both fundamental rotation frequencies and suggest an antisolar rotation profile. Next we use the missing-data multitaper power spectrum estimator to identify the oscillation modes responsible for the complex "scallop-shell" shape of the K2 light curve of EPIC 203354381. We argue that multitaper power spectrum estimators should be used for all time series with regular observing cadence.

79 ASTRONOMY AND ASTROPHYSICS↗

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

Model certainty in cellular network-driven processes with missing data

Mathematical models are often used to explore network-driven cellular processes from a systems perspective. However, a dearth of quantitative data suitable for model calibration leads to models with parameter unidentifiability and questionable predictive power. Here we introduce a combined Bayesian and Machine Learning Measurement Model approach to explore how quantitative and non-quantitative data constrain models of apoptosis execution within a missing data context. We find model prediction accuracy and certainty strongly depend on rigorous data-driven formulations of the measurement, and the size and make-up of the datasets. For instance, two orders of magnitude more ordinal (e.g., immunoblot) data are necessary to achieve accuracy comparable to quantitative (e.g., fluorescence) data for calibration of an apoptosis execution model. Notably, ordinal and nominal (e.g., cell fate observations) non-quantitative data synergize to reduce model uncertainty and improve accuracy. Finally, we demonstrate the potential of a data-driven Measurement Model approach to identify model features that could lead to informative experimental measurements and improve model predictive power.

59 BASIC BIOLOGICAL SCIENCES↗

Comparing Individualized Survival Predictions From Random Survival Forests and Multistate Models in the Presence of Missing Data: A Case Study of Patients With Oropharyngeal Cancer

Background: In recent years, interest in prognostic calculators for predicting patient health outcomes has grown with the popularity of personalized medicine. These calculators, which can inform treatment decisions, employ many different methods, each of which has advantages and disadvantages. Methods: We present a comparison of a multistate model (MSM) and a random survival forest (RSF) through a case study of prognostic predictions for patients with oropharyngeal squamous cell carcinoma. The MSM is highly structured and takes into account some aspects of the clinical context and knowledge about oropharyngeal cancer, while the RSF can be thought of as a black-box non-parametric approach. Key in this comparison are the high rate of missing values within these data and the different approaches used by the MSM and RSF to handle missingness. Results: We compare the accuracy (discrimination and calibration) of survival probabilities predicted by both approaches and use simulation studies to better understand how predictive accuracy is influenced by the approach to (1) handling missing data and (2) modeling structural/disease progression information present in the data. We conclude that both approaches have similar predictive accuracy, with a slight advantage going to the MSM. Conclusions: Although the MSM shows slightly better predictive ability than the RSF, consideration of other differences are key when selecting the best approach for addressing a specific research question. These key differences include the methods’ ability to incorporate domain knowledge, and their ability to handle missing data as well as their interpretability, and ease of implementation. Ultimately, selecting the statistical method that has the most potential to aid in clinical decisions requires thoughtful consideration of the specific goals.

60 APPLIED LIFE SCIENCES↗

Multitaper Magnitude‐Squared Coherence for Time Series With Missing Data: Understanding Oscillatory Processes Traced by Multiple Observables

To explore the hypothesis of a common source of variability in two time series, observers may estimate the magnitude-squared coherence (MSC), which is a frequency-domain view of the cross correlation. For time series that do not have uniform observing cadence, MSC can be estimated using Welch's overlapping segment averaging. However, multitaper has superior statistical properties to Welch's method in terms of the tradeoff between bias, variance, and bandwidth. The classical multitaper technique has recently been extended to accommodate time series with underlying uniform observing cadence from which some observations are missing. This situation is common for solar and geomagnetic data sets, which may have gaps due to breaks in satellite coverage, instrument downtime, or poor observing conditions. We demonstrate the scientific use of missing-data multitaper magnitude-squared coherence by detecting known solar mid-term oscillations in simultaneous, missing-data time series of solar Lyman α flux and geomagnetic Disturbance Storm Time index. Due to their superior statistical properties, we recommend that multitaper methods be used for all heliospheric time series with underlying uniform observing cadence.

Astro-statistics techniques (1886)↗

Spatio-Temporal Denoising Graph Autoencoders with Data Augmentation for Missing Photovoltaic Data Imputation

The integration of the global Photovoltaic (PV) market with real time data-loggers has enabled large scale PV data analytical pipelines for power forecasting and long-term reliability assessment of PV fleets. Nevertheless, the performance of PV data analysis heavily depends on the quality of PV timeseries data. This paper proposes a novel Spatio-Temporal Denoising Graph Autoencoder (STD-GAE) framework to impute missing PV Power Data. STDGAE exploits temporal correlation, spatial coherence, and value dependencies from domain knowledge to recover missing data. It is empowered by two modules. (1) To cope with sparse yet various scenarios of missing data, STD-GAE incorporates a domain-knowledge aware data augmentation module that creates plausible variations of missing data patterns. This generalizes STD-GAE to robust imputation over different seasons and environment. (2) STD-GAE nontrivially integrates spatiotemporal graph convolution layers (to recover local missing data by observed “neighboring” PV plants) and denoising autoencoder (to recover corrupted data from augmented counterpart) to improve the accuracy of imputation accuracy at PV fleet level. We have evaluated our proposed model on two realworld PV datasets. Experimental results show that STD-GAE can achieve a gain of 43.14% in imputation accuracy and remains less sensitive to missing rate, different seasons, and missing scenarios, compared with state-of-the-art data imputation methods such as MIDA and LRTC-TNN.

Fan, Yangxin↗

Daily, 30 m Resolution NDSI Data for the East River Watershed, CO for 2000-2020

This dataset contains daily Normalized Difference Snow Index (NDSI) values at 30 m spatial resolution for the East River watershed in Colorado, USA. The temporal range of these data includes water years 2001-2020. These data were created using the Spatial and Temporal Adaptive Reflectance Fusion Model (STARFM). This model fuses low spatial and high temporal resolution data from MODIS (500 m, daily) with high spatial and low temporal resolution data from Landsat (30 m, 16 days) to create a 30m synthetic daily snow product. This product allows for the analysis of historical snow covered area trends in the East River Watershed at fine spatiotemporal resolutions where it was not available previously. This research was performed as a part of the Department of Energy’s Subsurface Biogeochemical Research Program with the primary intent of better understanding the timing and spatial patterns of water delivery to the Critical Zone in mountain watersheds. Each .zip file contains one "water year" of data (October 1 - September 30; i.e., water year 2010 starts October 1, 2010 and ends September 30, 2011). Each zip file contains the following: STARFM daily Normalized Difference Snow Index (NDSI) fusion data files in GeoTiff format with one layer for each day between Landsat data acquisition dates (i.e., for dates of Landsat acquisition, the Landsat image is included for that date). The study area is located in an area of Landsat path overlap, so Landsat dates acquisitions are every 7-9 days. Landsat NDSI files containing the high spatial (30m), low temporal (7-9 days due to Landsat path overlap) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Dates for which no Landsat data were obtained are included as NoData layers. MODIS NDSI files containing the high temporal (daily), low spatial (500m) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Please note the MODIS data were resampled to 30m pixels for input into the STARFM model. The data have a scale factor of 10,000 and a no data value of -32767. The projection of all datasets is WGS 84 (EPSG: 4326), which has a latitude/longitude based degree resolution of 0.0002694946 X 0.0002694946, and approximates to the 30 m spatial resolution mentioned above. The Layer Index files in .csv format. They contain information for each layer in the above GeoTiff files regarding the corresponding date for each layer, the fraction of pixels in the image that contain valid data (missing data is due to either cloud cover or poor data quality; these values are not percent snow cover). Dates of Landsat overpass are indicated in these files. If no Landsat data were able to be obtained due to cloud cover or lack of Landsat Tier 1 data available on Google Earth Engine, this is also noted.

EARTH SCIENCE > CRYOSPHERE > SNOW/ICE↗

A General Spatiotemporal Imputation Framework for Missing Sensor Data

Many applications from precision agriculture, environmental monitoring and transportation networks rely on data collected across space and time over a large geographic area. Missing data poses a significant challenge for any data-driven inference and control tasks. Data imputation or the estimation of missing data can help fill these gaps by utilizing inherent spatial relationships and temporal patterns. A variety of spatiotemporal imputation models have been developed to address missing data in spatiotemporal datasets. However, these classical methods rely on the assumption that the underlying data follows a smooth trend and fail to provide accurate estimates when there is a large number of missing points in the data. Even though there are machine learning driven tensor completion approaches such as convolutional neural network based tensor completion (CoSTCo) that capture the non-linear relationships in the dataset, the transductive nature makes the algorithm less scalable. Thus, existing approaches for estimating the missing information do not effectively capture all dimensions of the spatiotemporal data structure, resulting in erroneous predictions and poor performance. The main contributions of this paper are: (1) We propose a novel inductive framework (G-LSTM) for missing data imputation that integrates a graph neural network with LSTMs to effectively capture both spatial and temporal dependencies. (2) Experimental results on a traffic dataset demonstrate that the proposed GNN integrated with an LSTM framework achieves improved imputation and maintains steady performance even when there are extreme missing conditions in comparison with the state-of-the-art imputation framework (i.e, CoSTCo). (3) The simulation results on a traffic network show up to 69% reduction in mean absolute error and 61% reduction in root mean square error when compared to CoSTCo.

Tharzeen, Aabila↗

Load Profile Inpainting for Missing Load Data Restoration and Baseline Estimation

This paper introduces a Generative Adversarial Nets (GAN) based, Load Profile Inpainting Network (Load-PIN) for restoring missing load data segments and estimating the baseline for a demand response event. The inputs are time series load data before and after the inpainting period together with explanatory variables (e.g., weather data). Here, we propose a Generator structure consisting of a coarse network and a fine-tuning network. The coarse network provides an initial estimation of the data segment in the inpainting period. The fine-tuning network consists of self-attention blocks and gated convolution layers for adjusting the initial estimations. Loss functions are specially designed for the fine-tuning and the discriminator networks to enhance both the point-to-point accuracy and realisticness of the results. We test the Load-PIN on three real-world data sets for two applications: patching missing data and deriving baselines of conservation voltage reduction (CVR) events. We benchmark the performance of Load-PIN with five existing deep-learning methods. Our simulation results show that, compared with the state-of-the-art methods, Load-PIN can handle varying-length missing data events and achieve 15-30% accuracy improvement.

14 SOLAR ENERGY↗

Assessing methods in fusion and fitting for time series construction in remote sensing-based earth observations

This study evaluates the comparative performance of spatiotemporal fusion and time-series fitting methods for constructing high-spatiotemporal-resolution remote sensing time-series data. Due to in-class similarity of fusion methods and fitting methods, we employ the Fit-FC (Fitting, spatial Filtering, and residual Compensation) model as a representative fusion method and the linear harmonic fitting model as a representative fitting method. Both Fit-FC and the linear harmonic fitting are widely used for high-spatiotemporal-resolution time-series data construction, and we modify the original Fit-FC model to enable automatic time-series fusion. To ensure data representativeness, we use 3 years (2019–2021) of Harmonized Landsat and Sentinel-2 surface reflectance datasets and Terra MCD43A4 products. Eight experimental regions are selected worldwide to guarantee generalization of the comparative performance between fusion and fitting methods, covering diverse land-use types (cropland, developed land, forest, and grassland) and varying climatological conditions. Time-series of NDVI and surface reflectance are analyzed under both actual observations and simulated data-missing scenarios. The constructed time-series data reveals that (1) the modified Fit-FC and linear harmonic fitting model achieve excellent performance in constructing high-resolution time-series images; (2) the fusion method outperforms the fitting method in constructing time-series of NDVI and surface reflectance images in cropland-, forest-, and grassland-dominated regions; (3) both methods achieve comparable performance in developed-dominated regions; (4) the fusion method is more robust to missing data, and better captures abrupt phenological transitions under conditions of continuous missing data; (5) the fitting method is computationally more efficient, making it suitable for large-scale time-series image reconstruction. This study provides valuable insights for selecting optimal strategies to generate high-resolution time-series images across diverse application scenarios and lays a foundation for extensions to other vegetation indices or land surface variables.

54 ENVIRONMENTAL SCIENCES↗

Accelerating full-waveform inversion using source stacking: synthetic experiments at the global scale in a realistic 3-D earth model

SUMMARY The spectral element method is currently the method of choice for computing accurate synthetic seismic wavefields in realistic 3-D earth models at the global scale. However, it requires significantly more computational time, compared to normal mode-based approximate methods. Source stacking, whereby multiple earthquake sources are aligned on their origin time and simultaneously triggered, can reduce the computational costs by several orders of magnitude. We present the results of synthetic tests performed on a realistic radially anisotropic 3-D model, slightly modified from model SEMUCB-WM1 with three component synthetic waveform ‘data’ for a duration of 10 000 s, and filtered at periods longer than 60 s, for a set of 273 events and 515 stations. We consider two definitions of the misfit function, one based on the stacked records at individual stations and another based on station-pair cross-correlations of the stacked records. The inverse step is performed using a Gauss–Newton approach where the gradient and Hessian are computed using normal mode perturbation theory. We investigate the retrieval of radially anisotropic long wavelength structure in the upper mantle in the depth range 100–800 km, after fixing the crust and uppermost mantle structure constrained by fundamental mode Love and Rayleigh wave dispersion data. The results show good performance using both definitions of the misfit function, even in the presence of realistic noise, with degraded amplitudes of lateral variations in the anisotropic parameter ξ. Interestingly, we show that we can retrieve the long wavelength structure in the upper mantle, when considering one or the other of three portions of the cross-correlation time series, corresponding to where we expect the energy from surface wave overtone, fundamental mode or a mixture of the two to be dominant, respectively. We also considered the issue of missing data, by randomly removing a successively larger proportion of the available synthetic data. We replace the missing data by synthetics computed in the current 3-D model using normal mode perturbation theory. The inversion results degrade with the proportion of missing data, especially for ξ, and we find that a data availability of 45 per cent or more leads to acceptable results. We also present a strategy for grouping events and stations to minimize the number of missing data in each group. This leads to an increased number of computations but can be significantly more efficient than conventional single-event-at-a-time inversion. We apply the grouping strategy to a real picking scenario, and show promising resolution capability despite the use of fewer waveforms and uneven ray path distribution. Source stacking approach can be used to rapidly obtain a starting 3-D model for more conventional full-waveform inversion at higher resolution, and to investigate assumptions made in the inversion, such as trade-offs between isotropic, anisotropic or anelastic structure, different model parametrizations or how crustal structure is accounted for.

Geochemistry & Geophysics↗

Explainable multi-fidelity Bayesian neural network for distribution system state estimation

Distribution System State Estimation (DSSE) is frequently constrained by limited real-time measurements, the uncertainties introduced by distributed energy resources, and the presence of bad data. To address them, this paper proposes an enhanced Multi-Fidelity Bayesian Neural Network (MFBNN) DSSE approach. A low-fidelity layer based on a Deep Neural Network (DNN) is first pre-trained on pseudo-measurement data to learn fundamental state features. Subsequently, a high-fidelity Bayesian Neural Network (BNN) layer leverages limited but high-quality real-time measurements to refine these features, thereby achieving accurate DSSE. Additionally, the deep SHapley Additive exPlanation (SHAP) is developed to quantify the influence of measurement data on DSSE through dual perspectives of global feature importance and local nodal contributions, establishing a hierarchical explainability framework for machine learning-based DSSE. Comparative studies conducted on the IEEE 13-bus system and a real-world 2135-node system from Dominion Energy demonstrate that the proposed method excels in estimation accuracy, even under situations of high noise levels, bad data, and missing data. Further comparisons with Weighted Least Squares (WLS) and other machine learning-based DSSE approaches verify that the proposed framework offers higher accuracy, improved interpretability, and enhanced robustness.

Bad data↗

BSEC flux towers: CSAT3B and TRH

The data were collected as part of the BSEC project, during the period from June 2025 to May 2026. Directory "broadway" contains data collected on a multi-level flux tower (US-BWf) in the Broadway East neighborhood (1808 North Patterson Park Ave., Baltimore City, MD 21213; LAT: 39o18'40.31'' N; LONG: 76o35'12.43'' W). At each of the four measurement heights (8.5 m, 11.1 m, 13.4 m, 15.9 m), a Campbell Scientific CSAT3B sonic anemometer was operated at 50 Hz to measure virtual temperature (tc) and three velocity components (u: 270 degrees; v: 180 degrees; w: vertical), and a RM Young temperature sensor (model 41382VC) was operated at 1 Hz inside a compact aspirated radiation shield (model 43502) to measure absolute temperature (T) and relative humidity (RH). Inside directory "broadway", directory "netcdf" contains data collected each day in 5-minute chunks that have been converted to NetCDF format (before quality checking), while "4hr" contains data arranged into 4-hour chunks (also in NetCDF format) that have been through basic quality checking steps (treating data points with nonzero diagnostic codes as missing data; fixing six or fewer consecutive missing data points using linear interpolation). Users are recommended to start with data in directory "4hr", while data in directory "netcdf" can be used for reference purposes.

Baltimore↗

LANL Meteorological Program: 2022 Data Completeness/Quality Report

Los Alamos National Laboratory (LANL) operates seven mesa-top instrumented meteorology towers: Technical Area (TA) 6, TA-49, TA-53, TA-54, TA-63, TA-54B, and TA-16B. An additional instrumented tower is located in Mortandad Canyon (TA-5 MDCN), and there is a rain gauge at North Community (NCOM), located within the town of Los Alamos. The 10 meter (m) towers at TA-63, TA- 54B, and TA-16B were installed in 2021, and will be included in the 2023 data completeness/quality report. A description of the meteorology monitoring network, prior to the installation of the TA-63, TA-54B, and TA-16B is found in Dewart and Boggs (2014). Four of the mesa-top towers (e.g., TA-6, TA-49, TA-53, and TA-54) are instrumented at the 1.2 m, 11.5 m, 23 m, and 46 m levels. In addition, the TA-6 tower is instrumented at the 92 m level. The TA-5 MDCN tower is 10 m in height and is instrumented at 1.2 m and 10 m. Data are collected and averaged every 15 minutes. Range checking is performed on each measurement every 15 minutes; data that are beyond normal ranges are eliminated from the data set and replaced by a code for missing data. In addition, data are reviewed weekly by qualified meteorologists to identify bad data not identified by the range checking technique. The data steward eliminates these data from the data set and replaces them with a code for missing data. The instrument technicians also review that data and schedule instrument replacement, as required. All instruments are calibrated at a frequency that meets the criteria identified in ANSI/ANS-3.11-2015. Data completeness is determined by the number of total 15-minute records available versus the number of possible measurements for the entire year. As a rule, the meteorologists do not attempt to estimate data that are eliminated as bad data. Original datalogger records, including bad data, can be recalled from program archival storage.

54 ENVIRONMENTAL SCIENCES↗

LANL Meteorological Program: 2023 Data Completeness/Quality Report

Los Alamos National Laboratory (LANL) operates seven mesa-top instrumented meteorology towers: Technical Area (TA) 6, TA-49, TA-53, TA-54, TA-63, TA-54B, and TA-16B. An additional instrumented tower is located in Mortandad Canyon (TA-5 MDCN), and there is a rain gauge at North Community (NCOM), located within the town of Los Alamos. The 10 meter (m) towers at TA-63, TA-54B, and TA-16B have been in testing since they were installed in 2021, and will be included in a future data completeness report. A description of the meteorology monitoring network, prior to the installation of the TA-63, TA-54B, and TA-16B is found in Dewart and Boggs (2014). Four of the mesa-top towers (e.g., TA-6, TA-49, TA-53, and TA-54) are instrumented at the 1.2 m, 11.5 m, 23 m, and 46 m levels. In addition, the TA-6 tower is instrumented at the 92 m level. The TA-5 MDCN tower is 10 m in height and is instrumented at 1.2 m and 10 m. Data are collected and averaged every 15 minutes. Range checking is performed on each measurement every 15 minutes; data that are beyond normal ranges are eliminated from the data set and replaced by a code for missing data. In addition, data are reviewed weekly by qualified meteorologists to identify bad data not identified by the range checking technique. The data steward eliminates these data from the data set and replaces them with a code for missing data. The instrument technicians also review that data and schedule instrument replacement, as required. All instruments are calibrated at a frequency that meets the criteria identified in ANSI/ANS-3.11-2015. Data completeness is determined by the number of total 15-minute records available versus the number of possible measurements for the entire year. As a rule, the meteorologists do not attempt to estimate data that are eliminated as bad data. Original datalogger records, including bad data, can be recalled from program archival storage.

54 ENVIRONMENTAL SCIENCES↗