Search NASA⌕ Search

SEARCH · Search NASA

Results for “Multivariate time series”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fast and Flexible Multivariate Time Series Subsequence Search

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which often contain several gigabytes of data. Surprisingly, research on MTS search is very limited. Most of the existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two algorithms to solve this problem (1) a List Based Search (LBS) algorithm which uses sorted lists for indexing, and (2) a R*-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences. Both algorithms guarantee that all matching patterns within the specified thresholds will be returned (no false dismissals). The very few false alarms can be removed by a post-processing step. Since our framework is also capable of Univariate Time-Series (UTS) subsequence search, we first demonstrate the efficiency of our algorithms on several UTS datasets previously used in the literature. We follow this up with experiments using two large MTS databases from the aviation domain, each containing several millions of observations. Both these tests show that our algorithms have very high prune rates (>99%) thus needing actual disk access for only less than 1% of the observations. To the best of our knowledge, MTS subsequence search has never been attempted on datasets of the size we have used in this paper.

Bhaduri, Kanishka↗

Comparison of Multivariate Time Series Prediction Techniques for Emulating Noah-LSM Soil Moisture Outputs

Land surface models are crucial tools for many earth science applications including numerical weather prediction, water resource and crop monitoring, and climatological analysis. Given a set of atmospheric forcings, seasonal data, and static parameters, models like Noah-LSM solve for land surface quantities including skin temperature, sensible heat flux, and soil moisture. While these calculations are theoretically robust, they are often computationally expensive. Since artificial neural networks (ANNs) are universal function approximators, they can learn to emulate the output of a deterministic numerical model given a time series of input forcings, with the learned ANN having substantially shorter execution time. The ANN could efficiently parameterize other models, generate ensembles, and provide first-guess inputs for retrievals. As such, with the goal of developing a model that efficiently mimics the output of Noah-LSM given NLDAS2 forcings on a region covering much of the central US, we examine and compare several neural network architectures for the multi-horizon multivariate time series forecasting problem. Recent literature includes a diverse set of approaches including autoregressive architectures like LSTM and GRU, parametric and non-parametric statistical predictors (ForecastNet and MQRNN), self-attention (LSTM-attention-LSTM), and temporal convovlution (DeepTCN). We implement several of these models for the Noah-LSM prediction task, highlighting the features and challenges for each and providing practical insight on the training process.

Mitchell Dodson↗

Capturing Multivariate Time Series Interactions to Detect High‑Risk Instability During Approach

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Capturing Complex Multivariate Time Series Interactions to Detect High-Risk Adverse Events During Flight

The reduction of aviation safety metrics below target thresholds continue to drive down the number of aviation fatalities and accidents. To meet future safety demands, sustained efforts by aviation agencies promoting safety assurance processes and systems have prompted ongoing research on identifying and mitigating in-flight risks. With the projected increase in passenger load factor and rollout of more autonomous systems into the national airspace, the need to detect high-risk events in-time or ahead-of-time is becoming increasingly crucial. New anomaly detection and precursor identification algorithms will need to scale to different airframes, levels of autonomy, and system complexity. While the pervasiveness of deep learning has resulted in the development of performant anomaly detection methods, these sophisticated models currently suffer from low end-user interpretability. Building off our previous work on identifying adverse events in multivariate flight data during descent, we propose a data-driven approach for detecting in-flight adverse events caused by the complex interplay of flight variables. Our approach utilizes ordinal patterns of important aircraft stability variables (e.g., airspeed and descent rate) to capture multivariate flight dynamics that can be used to predict the onset of unstable approaches, a high-risk adverse event that can occur during approach. Through the use of ordinal patterns, we aim to create more interpretable detection models of in-flight adverse events that can be translated to future autonomous systems without difficulty. Our analysis shows the presence of distinct ordinal pattern distributions that can be used to predict unstable approaches 1 minute ahead of time with an accuracy of 0.69 and a recall of 0.73 and 30 seconds ahead with an accuracy of 0.70 and a recall of 0.86.

Risk detection↗

Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations

This paper addresses the problem of detecting and characterizing local variability in time series and other forms of sequential data. The goal is to identify and characterize statistically significant variations, at the same time suppressing the inevitable corrupting observational errors. We present a simple nonparametric modeling technique and an algorithm implementing it-an improved and generalized version of Bayesian Blocks [Scargle 1998]-that finds the optimal segmentation of the data in the observation interval. The structure of the algorithm allows it to be used in either a real-time trigger mode, or a retrospective mode. Maximum likelihood or marginal posterior functions to measure model fitness are presented for events, binned counts, and measurements at arbitrary times with known error distributions. Problems addressed include those connected with data gaps, variable exposure, extension to piece- wise linear and piecewise exponential representations, multivariate time series data, analysis of variance, data on the circle, other data modes, and dispersed data. Simulations provide evidence that the detection efficiency for weak signals is close to a theoretical asymptotic limit derived by [Arias-Castro, Donoho and Huo 2003]. In the spirit of Reproducible Research [Donoho et al. (2008)] all of the code and data necessary to reproduce all of the figures in this paper are included as auxiliary material.

signal detection↗

Spatial estimation from remotely sensed data via empirical Bayes models

Multichannel satellite image data, available as LANDSAT imagery, are recorded as a multivariate time series (four channels, multiple passovers) in two spatial dimensions. The application of parametric empirical Bayes theory to classification of, and estimating the probability of, each crop type at each of a large number of pixels is considered. This theory involves both the probability distribution of imagery data, conditional on crop types, and the prior spatial distribution of crop types. For the latter Markov models indexed by estimable parameters are used. A broad outline of the general theory reveals several questions for further research. Some detailed results are given for the special case of two crop types when only a line transect is analyzed. Finally, the estimation of an underlying continuous process on the lattice is discussed which would be applicable to such quantities as crop yield.

Hill, J. R.↗

Biogeochemical Response to Mesoscale Physical Forcing in the California Current System

In the first part of the project, we investigated the local response of the coastal ocean ecosystems (changes in chlorophyll, concentration and chlorophyll, fluorescence quantum yield) to physical forcing by developing and deploying Autonomous Drifting Ocean Stations (ADOS) within several mesoscale features along the U.S. west coast. Also, we compared the temporal and spatial variability registered by sensors mounted in the drifters to that registered by the sensors mounted in the satellites in order to assess the scales of variability that are not resolved by the ocean color satellite. The second part of the project used the existing WOCE SVP Surface Lagrangian drifters to track individual water parcels through time. The individual drifter tracks were used to generate multivariate time series by interpolating/extracting the biological and physical data fields retrieved by remote sensors (ocean color, SST, wind speed and direction, wind stress curl, and sea level topography). The individual time series of the physical data (AVHRR, TOPEX, NCEP) were analyzed against the ocean color (SeaWiFS) time-series to determine the time scale of biological response to the physical forcing. The results from this part of the research is being used to compare the decorrelation scales of chlorophyll from a Lagrangian and Eulerian framework. The results from both parts of this research augmented the necessary time series data needed to investigate the interactions between the ocean mesoscale features, wind, and the biogeochemical processes. Using the historical Lagrangian data sets, we have completed a comparison of the decorrelation scales in both the Eulerian and Lagrangian reference frame for the SeaWiFS data set. We are continuing to investigate how these results might be used in objective mapping efforts.

Niiler, Pearn P.↗

Temporal Variability of Observed and Simulated Hyperspectral Earth Reflectance

The Climate Absolute Radiance and Refractivity Observatory (CLARREO) is a climate observation system designed to study Earth's climate variability with unprecedented absolute radiometric accuracy and SI traceability. Observation System Simulation Experiments (OSSEs) were developed using GCM output and MODTRAN to simulate CLARREO reflectance measurements during the 21st century as a design tool for the CLARREO hyperspectral shortwave imager. With OSSE simulations of hyperspectral reflectance, Feldman et al. [2011a,b] found that shortwave reflectance is able to detect changes in climate variables during the 21st century and improve time-to-detection compared to broadband measurements. The OSSE has been a powerful tool in the design of the CLARREO imager and for understanding the effect of climate change on the spectral variability of reflectance, but it is important to evaluate how well the OSSE simulates the Earth's present-day spectral variability. For this evaluation we have used hyperspectral reflectance measurements from the Scanning Imaging Absorption Spectrometer for Atmospheric Cartography (SCIAMACHY), a shortwave spectrometer that was operational between March 2002 and April 2012. To study the spectral variability of SCIAMACHY-measured and OSSE-simulated reflectance, we used principal component analysis (PCA), a spectral decomposition technique that identifies dominant modes of variability in a multivariate data set. Using quantitative comparisons of the OSSE and SCIAMACHY PCs, we have quantified how well the OSSE captures the spectral variability of Earth?s climate system at the beginning of the 21st century relative to SCIAMACHY measurements. These results showed that the OSSE and SCIAMACHY data sets share over 99% of their total variance in 2004. Using the PCs and the temporally distributed reflectance spectra projected onto the PCs (PC scores), we can study the temporal variability of the observed and simulated reflectance spectra. Multivariate time series analysis of the PC scores using techniques such as Singular Spectrum Analysis (SSA) and Multichannel SSA will provide information about the temporal variability of the dominant variables. Quantitative comparison techniques can evaluate how well the OSSE reproduces the temporal variability observed by SCIAMACHY spectral reflectance measurements during the first decade of the 21st century. PCA of OSSE-simulated reflectance can also be used to study how the dominant spectral variables change on centennial scales for forced and unforced climate change scenarios. To have confidence in OSSE predictions of the spectral variability of hyperspectral reflectance, it is first necessary for us to evaluate the degree to which the OSSE simulations are able to reproduce the Earth?s present-day spectral variability.

Roberts, Yolanda↗

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis↗

Anomaly Detection in Flight Operational Data Using Deep Learning

In this session, we demonstrate two recently developed deep learning models for anomaly detection in flight operational data by the Data Sciences Group at NASA Ames Research Center. The first model is Convolutional Variational Auto-Encoder (CVAE) [1], which is an unsupervised deep encoder-decoder model, designed specifically for finding anomalies in heterogeneous multivariate time series data. We will demonstrate its application to finding anomalies in streaming data from NASA’s Digital Information Platform’s Fuser source. CVAE identifies data instances that are not representative of expected nominal behavior as anomalous. Since it is an unsupervised approach, the flagged anomalies will need to be reviewed by the subject matter experts (SMEs) for validation and labeling and is designed to assist with vulnerability discovery within Safety Monitoring System programs. The second model is Robust and Explainable Semi-supervised Anomaly Detection (RESAD) model [2], which builds on CVAE to allow learning from both minimally labeled data (previously reviewed by the SMEs) as well as majority unlabeled data. RESAD takes advantage of graph theoretic techniques to propagate the labels from the labeled data to the unlabeled data based on a pre-defined similarity metric and structures the learned feature space from flight time-series so that data of the same class would cluster tightly together. This model characteristic is enabled by training with an augmented loss function and allows learning of a more informative feature space for down-stream tasks such as search and active learning. We demonstrate RESAD using data from the NASA DASHlink project [3].

anomaly detection↗

Fast Multivariate Search on Large Aviation Datasets

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which can contain up to several gigabytes of data. Surprisingly, research on MTS search is very limited. Most existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two provably correct algorithms to solve this problem (1) an R-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences, and (2) a List Based Search (LBS) algorithm which uses sorted lists for indexing. We demonstrate the performance of these algorithms using two large MTS databases from the aviation domain, each containing several millions of observations Both these tests show that our algorithms have very high prune rates (>95%) thus needing actual

Bhaduri, Kanishka↗

Investigating Stable Oxygen and Carbon Isotopic Variability in Speleothem Records Over the Last Millennium Using Multiple Isotope-Enabled Climate Models

The incorporation of water isotopologues into the hydrology of general circulation models (GCMs) facilitates the comparison between modeled and measured proxy data in paleoclimate archives. However, the variability and drivers of measured and modeled water isotopologues, as well as the diversity of their representation in different models, are not well constrained. Improving our understanding of this variability in past and present climates will help to better constrain future climate change projections and decrease their range of uncertainty. Speleothems are a precisely datable terrestrial paleoclimate archives and provide well-preserved (semi-)continuous multivariate isotope time series in the lower latitudes and mid-latitudes and are therefore well suited to assess climate and isotope variability on decadal and longer timescales. However, the relationships of speleothem oxygen and carbon isotopes to climate variables are influenced by site-specific parameters, and their comparison to GCMs is not always straightforward. Here we compare speleothem oxygen and carbon isotopic signatures from the Speleothem Isotopes Synthesis and Analysis database version 2 (SISALv2) to the output of five different water-isotope-enabled GCMs (ECHAM5-wiso, GISS-E2-R, iCESM, iHadCM3, and isoGSM) over the last millennium (850–1850 CE). We systematically evaluate differences and commonalities between the standardized model simulation outputs. The goal is to distinguish climatic drivers of variability for modeled isotopes and compare them to those of measured isotopes. We find strong regional differences in the oxygen isotope signatures between models that can partly be attributed to differences in modeled surface temperature. At low latitudes, precipitation amount is the dominant driver for stable water isotope variability; however, at cave locations the agreement between modeled temperature variability is higher than for precipitation variability. While modeled isotopic signatures at cave locations exhibited extreme events coinciding with changes in volcanic and solar forcing, such fingerprints are not apparent in the speleothem isotopes. This may be attributed to the lower temporal resolution of speleothem records compared to the events that are to be detected. Using spectral analysis, we can show that all models underestimate decadal and longer variability compared to speleothems (albeit to varying extents). We found that no model excels in all analyzed comparisons, although some perform better than the others in either mean or variability. Therefore, we advise a multi-model approach whenever comparing proxy data to modeled data. Considering karst and cave internal processes, e.g., through isotope-enabled karst models, may alter the variability in speleothem isotopes and play an important role in determining the most appropriate model. By exploring new ways of analyzing the relationship between the oxygen and carbon isotopes, their variability, and co-variability across timescales, we provide methods that may serve as a baseline for future studies with different models using, e.g., different isotopes, different climate archives, or different time periods.

isotope-enabled simulations↗

The Statistical Consulting Center for Astronomy (SCCA)

The process by which raw astronomical data acquisition is transformed into scientifically meaningful results and interpretation typically involves many statistical steps. Traditional astronomy limits itself to a narrow range of old and familiar statistical methods: means and standard deviations; least-squares methods like chi(sup 2) minimization; and simple nonparametric procedures such as the Kolmogorov-Smirnov tests. These tools are often inadequate for the complex problems and datasets under investigations, and recent years have witnessed an increased usage of maximum-likelihood, survival analysis, multivariate analysis, wavelet and advanced time-series methods. The Statistical Consulting Center for Astronomy (SCCA) assisted astronomers with the use of sophisticated tools, and to match these tools with specific problems. The SCCA operated with two professors of statistics and a professor of astronomy working together. Questions were received by e-mail, and were discussed in detail with the questioner. Summaries of those questions and answers leading to new approaches were posted on the Web (www.state.psu.edu/ mga/SCCA). In addition to serving individual astronomers, the SCCA established a Web site for general use that provides hypertext links to selected on-line public-domain statistical software and services. The StatCodes site (www.astro.psu.edu/statcodes) provides over 200 links in the areas of: Bayesian statistics; censored and truncated data; correlation and regression, density estimation and smoothing, general statistics packages and information; image analysis; interactive Web tools; multivariate analysis; multivariate clustering and classification; nonparametric analysis; software written by astronomers; spatial statistics; statistical distributions; time series analysis; and visualization tools. StatCodes has received a remarkable high and constant hit rate of 250 hits/week (over 10,000/year) since its inception in mid-1997. It is of interest to scientists both within and outside of astronomy. The most popular sections are multivariate techniques, image analysis, and time series analysis. Hundreds of copies of the ASURV, SLOPES and CENS-TAU codes developed by SCCA scientists were also downloaded from the StatCodes site. In addition to formal SCCA duties, SCCA scientists continued a variety of related activities in astrostatistics, including refereeing of statistically oriented papers submitted to the Astrophysical Journal, talks in meetings including Feigelson's talk to science journalists entitled "The reemergence of astrostatistics" at the American Association for the Advancement of Science meeting, and published papers of astrostatistical content.

Akritas, Michael↗

Identification of multivariable high performance turbofan engine dynamics from closed loop data

The multivariable instrumental variable/approximate maximum likelihood (IV/AML) method or recursive time-series analysis is used to identify the multivariable (four inputs-three outputs) dynamics of the Pratt and Whitney F100 engine. A detailed nonlinear engine simulation is used to determine linear engine model structures and parameters at an operating point using open loop data. Also, the IV/AML method is used in a direct identification mode to identify models from actual closed loop engine test data. Models identified from simulated and test data are compared to determine a final model structure and parameterization that can predict engine response for a wide class of inputs. The ability of the IV/AML algorithm to identify useful dynamic models from engine test data is assessed.

Merrill, W.↗