Search NASA⌕ Search

SEARCH · Search NASA

Results for “Time series data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Large Deviations Anomaly Detection (LAD) for collection of multivariate time series data: Applications to COVID-19 data

Time series anomaly detection is frequently used to identify extreme behaviors within a single time series. Identifying extreme trends in relation to a collection of other time series, on the other hand, is frequently of significant interest, such as in public health policy, social justice, and pandemic propagation. Using concepts from large deviations theory , we propose an algorithm that can scale to large collections of time series data. This paper expands on the LAD algorithm presented in Guggilam et al. (2022). The proposed algorithm is an online anomaly detection method for identifying anomalies in a collection of multivariate time series that takes advantage of the algorithm’s ability to scale to high-dimensional data. We show how the proposed Large Deviations Anomaly Detection (LAD) algorithm can be used to identify regions with anomalous trends in COVID-19 cases, deaths, biweekly growth rates, vaccinations, and fatality rates. Several of the observed anomalous trends are associated with regions that have demonstrated poor response to the COVID pandemic.

97 MATHEMATICS AND COMPUTING↗

Elastic Changepoint Detection for Globally-indexed Functional Time Series Data with Climate Applications

Changepoint detection is a vital tool in the application of climate data analysis. Numerous types of climate observation data are most properly represented by functional time series, implying a need for accurate changepoint detection methods applicable to functional time series data. Such data taken at a global scale often contain both spatial heterogeneity and dependence as well as phase (time) misalignment. In this report, we present methods which can detect spatially-dependent changepoints while allowing different estimates of change time and change strength depending on location. Additionally, we provide extensions to this spatially-predicted model which controls for phase variability among observations. Our methods provide the ability to detect a single change, or control for epidemic changes (where a “return-to-normal” change is more likely to be detected than the initial change). We showcase results analyzing the June 1991 eruption of Mt. Pinatubo, where our methods demonstrate the ability to accurately detect both single and epidemic changepoints even in the presence of strong seasonal variability. We find that our spatially-predicted model improves the detection of relevant changepoints versus methods which do not take spatial information into account, and we find that controlling for phase variability helps to control the false discovery rate during the detection process.

54 ENVIRONMENTAL SCIENCES↗

Leveraging Gaussian Mixture Models for Detecting Anomalies in Time-Series Data

Test systems must be capable of classifying measured data as expected or anomalous in real time. Anomalous results may portend system failure, and, if undetected, may result in damage to the unit, test equipment, or potential harm to personnel. This report investigates the use of Gaussian Mixture Models (GMMs) as a clustering tool in classifying time-series data.

Wilke, Rudeger H.T. [Sandia National Laboratories ↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Aboveground Biomass Estimation Using NISAR Simulated ALOS-2 Time Series Data

Aboveground biomass (AGB) is a critical parameter to better understand the global carbon cycle and to develop sustainable forest management. However, a large uncertainty prevails. L-band SAR data have demonstrated strong potential to accurately retrieve AGB over low-biomass regions (<100 Mg ha-1). The upcoming NASA-ISRO Synthetic Aperture Radar mission will collect data at L- and S-band over earth’s landmass with a repeat period of 12 days, allowing us to have ample data for monitoring biomass and its dynamics. One of the key science requirements of the mission is to produce annual AGB maps at 1-ha resolution with RMS accuracy of 20 Mg/ha for 80 percentage of area over low-biomass regions in Calibration/Validation sites. The NISAR biomass algorithm will generate AGB maps based on the parameterization of semi-empirical model along with NISAR time-series dual pol data (HH and HV). To calibrate and validate the model for mission requirements, the mission will use reference estimates of AGB produced from ground inventory plots and airborne LiDAR data collected over selected sites distributed across different global ecoregions. This paper presents the initial results of the calibration/validation of the NISAR AGB retrieval algorithm over the Lenoir Landing (LENO), Alabama, USA site using NISAR simulated ALOS-2 time series data. Five multi-temporal dual-pol HH and HV NISAR Simulated ALOS 2 data collections were used as input to assess the performance of the model. The model AGB retrieval results shows that the NISAR model was able to achieve RMS accuracy within 20 Mg/ha.

Ramachandran, Naveen [Jet Propulsion Laboratory, C↗

A Two-Step Time-Series Data Clustering Method for Building-Level Load Profile

Residential and commercial buildings have huge potential to contribute value to improve grid resilience by participating grid services. To reveal the significant value, it is critical to estimate the grid service capability from these buildings. Unlike the large-scale distributed energy resources such as wind and solar farms, those buildings need to participate grid services in aggregation, not by individual. Therefore, it is important to appropriately group buildings for aggregation. The load profiles in the same group will have similar characteristics at the same time step, so grid operators can send the grid service signal to the customer group with a higher chance to respond at that time step. In this paper, we develop a load profile clustering method to classify the building-level load profiles for grid service capability estimation. In our two-step clustering approach, we first calculate the total load consumption for each building, clustering the load profiles based on energy consumption level. Then, we further cluster the load profiles in each energy cluster based on the load shape. The parameter selection for each clustering step is discussed. The proposed method is applied on actual building-level load profiles, and the results have proved the effectiveness of this method.

advanced metering infrastructure (AMI)↗

A Two-Step Time-Series Data Clustering Method for Building-Level Load Profile

Residential and commercial buildings have huge potential to contribute value to improve grid resilience by participating grid services. To reveal the significant value, it is critical to estimate the grid service capability from these buildings. Unlike the large-scale distributed energy resources such as wind and solar farms, those buildings need to participate grid services in aggregation, not by individual. Therefore, it is important to appropriately group buildings for aggregation. The load profiles in the same group will have similar characteristics at the same time step, so grid operators can send the grid service signal to the customer group with a higher chance to respond at that time step. In this paper, we develop a load profile clustering method to classify the building-level load profiles for grid service capability estimation and the results have proved the effectiveness of this method.

AMI↗

Extreme Weather Events and the Impact on PV Time Series Data

The impact of extreme weather events on PV performance was studied by comparing the National Oceanic and Atmospheric Administration database on severe weather with the PV Fleet database on continuous PV performance. We identified 170 systems that were immediately impacted by weather events. These severe weather events lead to a median loss of only 1% of annual production. However, flooding and high wind events were found to have an extremely long tail extending to 60 % loss showing that these discrete events can pose a substantial risk to PV systems. Besides the short-term impact of lost production due to outages, we also found a statistically significant increased performance loss rate (PLR) for high wind events comparing PLR before and after these weather events. In addition, hail events caused a higher PLR for 2 out of 3 systems. More data are required to better quantify the impact, but these first results illustrate the substantial risk these events pose short-and long-term.

degradation↗

A Two-Step Time-Series Data Clustering Method for Building-Level Load Profile: Preprint

Residential and commercial buildings have huge potential to contribute value to improve grid resilience by participating grid services. To reveal the significant value, it is critical to estimate the grid service capability from these buildings. Unlike the large-scale distributed energy resources such as wind and solar farms, those buildings need to participate grid services in aggregation, not by individual. Therefore, it is important to appropriately group buildings for aggregation. In this paper, we develop a load profile clustering method to classify the building-level load profiles for grid service capability estimation. In our two-step clustering approach, we first calculate the total load consumption for each building, clustering the load profiles based on energy consumption level. Then, we further cluster the load profiles in each energy cluster based on the load shape. The parameter selection for each clustering step is discussed. The proposed method is applied on actual building-level load profiles, and the results have proved the effectiveness of this method.

advanced metering infrastructure (AMI)↗

Generating Synthetic Time Series Photovoltaic Data with Real-World Physical Challenges and Noise for Use in Algorithm Test and Validation

The PV Fleet Data Initiative and other projects seek the develop algorithms for automated analysis of PV time series data for extraction of statistical information and other parameters of the data such as degradation rates, soiling loss information, tracker performance, clipping or curtailment, system availability and other valuable information. While there is a vast body of PV data available for application of said extraction algorithms it is difficult to validate these algorithms because the true parameters to be extracted are not known. There has been a wide use of synthetic data in the literature for algorithm validation but this synthetic data is typically very bounded by the problem or topic at hand. The PV Fleet Data Initiative project has demonstrated that real time series PV data almost always includes a host of data quality and physical problems that, in reality, any automated PV abstraction algorithm must handle appropriately. For this reason, this work describes the development of a complex synthetic PV times series data set that includes data quality and physical problems that have been experienced in real world PV data. The various quality and physical problems are documented in the synthetic data so that users can test the validity of various PV extraction algorithms as well as develop new algorithms to solve problems this data set can support.

14 SOLAR ENERGY↗

Schneider Springs Fire Study 2023 for Ecosystem Respiration Rates: Surface Water Chemistry and Hydrologic Sensor Data across the Yakima River Basin, Washington, USA (v2)

This dataset supports a broader study examining the drivers of spatial variability in wildfire impacts across the Yakima River Basin. Data provided within this dataset were generated from sample collection across 17 total sites (8 sites affected by a recent wildfire, 9 sites unaffected by a recent wildfire) within multiple rivers throughout the Yakima River Basin in Washington, USA from May-July 2023. Fire affected sites are defined as those affected by the 2021 Schneider Springs Fire, based on the drainage area of the streams being within the 2021 Schneider Springs Fire burn perimeter or not (Figure 1, below). The contents include surface water geochemistry data (dissolved organic carbon; total dissolved nitrogen; total suspended solids); short-term sonde data (specific conductivity; turbidity; pH; chlorophyll A; temperature); stream depth data; stream velocity; manual chamber open channel respiration data; sensor time-series data (oxygen; water pressure; barometric pressure); field metadata (including qualitative information on in stream and river corridor characteristics); and environmental context photos taken in the field. The dataset also includes a summary file of the sensor data and plots of the sensor data. Sensors were only recovered at 15 out of the 17 sites, and not all sensors were recovered at all 15 sites (see Methods section for more details), therefore all data does not exist at all sites. Data from a 2022 study at the same sites, as well as additional sites, can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566. The data package was originally published in November 2023. It was updated in June 2025 (v2; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder with field photos and one main data folder with two subfolders. The main data folder consists of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) field protocol; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) stream depth and averages. The sensor data subfolder consists of (1) sensor installation methods summary; (2) stream velocity; and (3) six subfolders. The BarotrollAtm (barometric pressure; temperature), DepthHOBO (water pressure; temperature), MantaRiver (specific conductivity; turbidity; pH; chlorophyll A; temperature), EXO (specific conductivity; pH; temperature), miniDOT (dissolved oxygen; temperature), and miniDOTManualChamber (dissolved oxygen; temperature) contain time-series data, plots, and summary files. The sample data subfolder consists of (1) total suspended solids (TSS) data; (2) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (3) total dissolved nitrogen (TN) data and averages; and (4) methods codes. All files are .csv, .pdf, .jpg, .jpeg, or .mov.

54 ENVIRONMENTAL SCIENCES↗

Mixed Delay/Nondelay Embeddings Based Neuromorphic Computing with Patterned Nanomagnet Arrays

Patterned nanomagnet arrays (PNAs) have been shown to exhibit a strong geometrically frustrated dipole interaction. Some PNAs have also shown emergent domain wall dynamics. Previous works have demonstrated methods to physically probe these magnetization dynamics of PNAs to realize neuromorphic reservoir systems that exhibit chaotic dynamical behavior and high-dimensional nonlinearity. These PNA reservoir systems from prior works leverage echo state properties and linear/nonlinear short-term memory of component reservoir nodes to map and preserve the dynamical information of the input time-series data into nondelay spatial embeddings. Such mappings enable these PNA reservoir systems to imitate and predict/forecast the input time series data. However, these prior PNA reservoir systems are based solely on the nondelay spatial embeddings obtained at component reservoir nodes. As a result, they require a massive number of component reservoir nodes, or a very large spatial embedding (i.e., high-dimensional spatial embedding) per reservoir node, or both, to achieve acceptable imitation and prediction accuracy. These requirements reduce the practical feasibility of such PNA reservoir systems. To address this shortcoming, we present a mixed delay/nondelay embeddings-based PNA reservoir system. Our system uses a single PNA reservoir node with the ability to obtain a mixture of delay/nondelay embeddings of the dynamical information of the time-series data applied at the input of a single PNA reservoir node. Our analysis shows that when these mixed delay/nondelay embeddings are used to train a perceptron at the output layer, our reservoir system outperforms existing PNA-based reservoir systems for the imitation of NARMA 2, NARMA 5, NARMA 7, and NARMA 10 time series data, and for the short-term and long-term prediction of the Mackey Glass time series data.

Ti, Changpeng↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗