Search NASA⌕ Search

SEARCH · Search NASA

Results for “time-series analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Full-polarization millimeter wavelength variability of Sagittarius A * during the 2018 EHT campaign

Context. Sagittarius A* (Sgr A*), the supermassive black hole at the center of the Milky Way, provides a unique laboratory to study accretion dynamics and plasma processes near the event horizon. Aims. We investigated the variability and polarization properties of Sgr A* using ALMA observations during the 2018 Event Horizon Telescope campaign. Methods. We analyzed high-cadence full-polarization light curves from ALMA at millimeter wavelengths, performed time-series analysis, and investigated the temporal behavior during an X-ray flare observed by Chandra on 2018 April 24. The variability characteristics are compared with expectations from standard accretion flow models. Results. We find low variability in total intensity (σ/μ < 10%), but significantly higher variability in linear and circular polarization (∼30% and ∼50%, respectively). A time-series analysis reveals red-noise variability, with power spectral densities between −2 and −3 across all Stokes parameters. Polarized intensity shows stable intra-day timescales, while total intensity exhibits more variable timescales, suggesting distinct emission regions, with polarization likely arising from a coherent structure. On April 24, a statistically significant inter-band delay in polarized intensity coincides with a near-simultaneous X-ray and millimeter peak that deviates from the typical delayed flare scenario. This event also features enhanced millimeter variability and coherent polarization loop evolution. The observed simultaneity challenges standard models of transient synchrotron emission with cooling delays, favoring instead a scenario of continuous energy injection in an optically thin region. Conclusions. Our results offer new constraints on the physical mechanisms driving variability in Sgr A*, and provide key observational input for refining theoretical models of accretion and plasma behavior in the vicinity of supermassive black holes.

Galaxy: center↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Range Hood Use and Effectiveness in Reducing Indoor Air Pollution During Gas and Induction Cooking

The Cooking Energy and Ventilation Impacts on Children's Asthma (CEVICA) study measured cooking frequency, range hood use, indoor air quality and respiratory health indicators of children with asthma living in homes with gas stoves in California's San Joaquin Valley. The study installed electric induction stoves and repeated measurements over three 2-week intensive periods, at baseline and at the end of two consecutive 3-month study phases. Stove replacements occurred at the start of Phase 1 or Phase 2 by random assignment. There were 4184 cooking events identified by automated analysis of time-series data from temperature sensors mounted above the cooktops and 1038 related range hood usage events detected from data recorded by anemometers, smart plugs, or motor loggers. Analysis of 1-minute resolved PM2.5 and NO 2 data identified and quantified 2685 PM 2.5 events and 2606 NO 2 events. Range hood use was characterized as a binary variable (>3 min vs. <3 min use). Range hood use was more common during cooking events associated with particle emissions and longer cooking durations. PM 2.5 concentrations during events with range hood use were comparable to those without use, which could result from limited effectiveness or if range hoods were preferentially used during higher-emission cooking scenarios. In homes with gas cooking, integrated NO 2 concentrations were about 45 percent higher during cooking events with no range hood use compared to those range hood use. The lowest pollutant levels were observed when the range hood operated for more than half of the cooking duration. These findings show that operation of venting range hood during cooking can substantially reduce short-term indoor exposure NO 2 in homes with gas cooking.

Fang, Yi↗

Rapid organic carbon spiraling in a headwater stream linked with streamflow, biogeochemistry, and canopy phenology

Headwater streams are abundant worldwide and important to global biogeochemical cycles, serving as critical processors and transporters of C. C spiraling is a useful way to understand the retention and mineralization of organic C (OC) in streams. However, analyses of seasonal and interannual variability in OC spiraling are currently limited. In this study, we aimed to understand the temporal patterns and driving mechanisms of OC spiraling, which will inform our understanding of future OC changes under climate change. We used 7 y of daily data in a small headwater stream (Walker Branch, Tennessee, USA) to assess seasonal and interannual variability in OC spiraling length (S OC ) and mineralization velocity (v fOC ), as well as their potential related variables. On average, S OC in Walker Branch was ~10× shorter than in previously studied small streams, indicating strong connections between the water column and the benthic environment where OC mineralization mostly takes place. OC spiraling was faster during the more biologically active periods of spring and autumn compared with more elongated OC spiraling in summer and winter, when OC retention was lower and downstream transport was higher. Gross primary production (GPP) was most strongly related to S OC and v fOC . Photosynthetically active radiation (PAR) and NO 3 − were also positively and negatively related to v fOC , respectively. Trends toward earlier and longer canopy cover and reduced GPP and PAR may result in longer S OC and slower v fOC , reducing localized instream processing of OC and potentially shunting more OC downstream. However, long-term observations indicate reduced NO 3 − at Walker Branch, suggesting opposing effects to those of GPP and PAR, leading to faster v fOC and greater OC retention. Time-series analyses of OC spiraling in streams can enhance our understanding of current and future responses of OC processing and downstream transport to climate change, as well as implications for downstream OC dynamics.

biological activity↗

EV Profile Capture

NextGen Profiles' EV profile capture efforts aimed to explore the variance in performance and evaluate how different operational conditions influence production EV charging behavior. Data were collected at a frequency of 10 Hz from both the EV and EVSE during each charge session. These charge session parameters were then entered into a time-series database for further analysis. The data were gathered under different operational conditions to examine the effects of various factors such as battery state of charge, battery temperature, vehicle condition, smart charge management, and EVSE limitations. The EV profile capture dataset includes extensive high-power charging data from 16 different EVs—comprising light-, medium-, and heavy-duty vehicles—along with EVSE from various suppliers. To protect confidentiality, the EV and EVSE metadata are anonymized, and the publicly released datasets are aggregated to 0.1-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

EVSE Characterization

NextGen Profiles' EVSE characterization efforts explored performance variability in production EVSE through the use of EV emulation equipment and assessed how different operational conditions influence charging behavior. Data were collected at a frequency of 10 Hz from both the EV emulator and EVSE during each charge session and stored in a time-series database for further analysis. As part of the NextGen Profiles project, characterization of high-power EVSE was performed on both conductive and wireless charging infrastructure; however, only conductive charging data are currently included in this repository. This EVSE characterization was performed over a range of DC output currents and voltages, covering both nominal and off-nominal test conditions. This EVSE characterization dataset includes high-power charging data from two types of 350-kW-capable EVSE using liquid-cooled Combined Charging System-1 (CCS1, North American version) cables and connectors. To protect confidentiality, all EVSE metadata are anonymized, and the publicly released datasets are metered at 10-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Three-Dimensional Grid Visualization for Planning Activities: A Dubai Case Study

National Laboratory of the Rockies (NLR), in collaboration with the Dubai Electricity and Water Authority (DEWA) and Infra-X, has undertaken the Energy Visualization Analysis Project. The aim of this project is to enhance analytical and 3D visualization capabilities for distribution network planning and renewable energy integration. As modern grid continues to evolve with large-scale solar PV deployment and emerging distributed energy resources (DERs), the ability to effectively analyze, visualize, and communicate complex grid behaviors has become increasingly critical. The project focuses on developing empirical use cases based on real distribution feeder data and engineering workflows, ensuring the outcomes are directly aligned with operational environment. Through time-series power flow simulations and nodal hosting capacity analysis, the study quantifies the impacts of high PV penetration on voltage and thermal limits within representative 11 kV feeders. These analyses identify specific nodes and conditions where DER integration challenges arise. Furthermore, a Battery Energy Storage System (BESS) optimization algorithm was applied to determine the optimal size and placement of storage systems that can mitigate network constraints and enhance hosting capacity. The comparative results between base-case and BESS-augmented scenarios clearly demonstrate improvements in network stability and load management efficiency. In parallel, the NLR team developed an immersive 3D visualization framework, enabling interactive exploration of grid simulations using commodity head-mounted display (HMD) systems. This framework transforms conventional 2D simulation data into spatially intuitive visual environments - allowing engineers to analyze feeder conditions, PV hosting potential, and BESS effects in real time. This report represents the first foundational phase in establishing a visualization-driven analytical ecosystem. It provides a methodological foundation for data integration, visualization architecture, and simulation-based decision support, paving the way for large-scale adoption of immersive visualization across DEWA's Smart Grid Initiative, R&D activities, and future network resilience studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Measurement-informed Dynamic Aggregation of Distribution Systems

This paper proposes a measurement-informed dynamic aggregation methodology in order to create equivalent representations of distribution systems that are compatible with large-scale transmission analysis. By optimizing an equivalent feeder parameters using time-series measurements of active power, reactive power, and voltage at the Point of Interconnection (POI), the approach yields simplified yet dynamically accurate equivalents. Implemented in PSCAD with models of photovoltaic–battery systems, three-phase motors, and static loads, the method employs hybrid differential evolution and bounded least-squares optimization laying the foundation for for real-time state estimation and optimized sensor placement in distribution networks.

Ahmed, Kazi Ishrak [University of Tennessee, Knoxv↗

Experimental Investigation of Subcritical Neutron and Gamma Noise Methods at the Seven Percent Critical Experiment (7uPCX)

As part of a collaborative international effort organized by the Lawrence Livermore National Laboratory (LLNL), with key participants from the Institut de radioprotection et de sûreté nucléaire (IRSN), Los Alamos National Laboratory (LANL), and the Sandia National Laboratories (SNL), a series of high-multiplication subcritical neutron and gamma noise measurements was planned and executed. The primary aim of this article was to advance detector technology, assess the validity of gamma noise for subcriticality measurements, and nuclear criticality safety, focusing on collecting list-mode or time-series data from various reactor configurations with multiplication values ranging from 20 to 310. This comprehensive dataset enabled a detailed comparative analysis of multiple detector systems and the results of both neutron and gamma noise measurements. In this work, we focus on experimentally comparing the results from neutron and gamma noise measurements. We note good agreement between estimations of the prompt neutron decay constant and demonstrate the effects of changing reactor geometry on the efficiency of the differing methods. Our results agree well with independent experimental measurements and simulations performed by LLNL.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Single nucleotide variants drive evolutionary phage-host arms race in anaerobic carbon dioxide-converting microbiome

Microbial bioconversions are shaped by environmental perturbations and the adaptation of resident microbiomes. Prokaryotes coexist with bacteriophages, yet their coevolutionary trajectories remain underexplored. Here, we investigate the effects of a cultivation vessel leak on an anaerobic consortium performing carbon dioxide reduction. Using time-series shotgun metagenomic sequencing, we reconstruct microbial and viral genomes to track community shifts. We further apply single-nucleotide variant profiling and CRISPR array analysis to monitor viral microdiversity and host defense mechanisms. After bioaugmentation restores bioconversion efficiency, the consortium undergoes pronounced restructuring, with new dominant taxa emerging from the rare biosphere. We identify patterns consistent with phage predation selectively removing certain species, while others exhibit resilience to infection. This shift aligns with a widespread viral outbreak and a transient increased frequency of single nucleotide variants in bacterial CRISPR–Cas defense genes. Expansion of CRISPR spacers further supports that CRISPR-mediated processes influence microbial resilience. Concurrently, phages infecting resilient hosts exhibited adaptive evolution, marked by high genetic heterogeneity. Selective pressure varies across their genomes, targeting infectivity genes and protospacer-adjacent motifs. These findings highlight a dynamic evolutionary arms race driven by the selection of beneficial genetic variants, providing a mechanistic framework for multi-omics investigations, and informing biotechnological applications, including phage-based microbiome manipulation.

Ghiotto, G↗

AI-Batt (Autonomous Identification of Battery Life Models) [SWR 21-36]

Autonomous Identification of Battery Life Models (AI-Batt) AI-Batt is a MATLAB code base for developing lifetime models for batteries from accelerated aging data. The code base provides many functions for processing, visualizing, and modeling battery aging data, making the data processing, exploration, and modeling workflow substantially faster. These tools are tailored for working with battery aging data sets, which usually consist of many separate time-series for each cell, with many test conditions and possible replicates at each condition, which makes it difficult to simply process or visualize the data set. Complex modeling tasks, such as cross-validation, sensitivity analysis, and uncertainty quantification have been implemented to enable thorough statistical investigation of model predictions. Additionally, several machine-learning algorithms are implemented to autonomously identify suitable models via symbolic regression. Data processing functions automatically cast data from the struct data type, which is commonly used to store experimental data, but is not an acceptable input for most algorithms, to the table data type, which can be easily used as input to any optimization algorithm. Also, the data can be separated into time-invariant and time-variant data tables, which is helpful for exploring the data set as well as developing separate models for time-variant and time-invariant aging mechanisms. For example, in aging tests with constant temperature, temperature is a time-invariant experimental condition. Visualization tools enable plotting of data, model fits, and model simulations possible with single-line function calls, empowering data exploration of complex data sets with both time-varying and time-invariant trends. Plots can be automatically generated for the whole data set, or separated by data group (groups of test replicates) or individual data series. Data points or data series can be automatically colored by the value of a variable with a variety of color maps, and model predictions can also be colored by the value of a fit statistic. Comparisons between data sets and the predictions/simulations of different models on the same data set can be easily plotted as well. Distributions of parameter values from bootstrap resampling can be plotted to visualize the reliability of parameter estimation, or determine any correlations between parameters. Modeling tools handle the complex task of creating and parsing symbolic equations for modeling battery lifetime. Equations are parsed to grab relevant data variables, parameter values, or specified sub-models for input into optimization, evaluation, or simulation functions. Models can be optimized locally (one set of parameters for each data series), bi-level (some parameters shared across the data set), or globally (single set of parameters for all data). Functions implementing symbolic regression algorithms help users to discover effective model equations, even in poorly sampled, high-dimensional data.

Smith, Kandler [National Renewable Energy Lab. (NR↗

The U.S. Agrivoltaic Shading Tool: A National-Scale Interface for Modeling Light and Shade Patterns in Ten Common Agrivoltaic Configurations

Agrivoltaic systems are dual-use configurations that co-locate agriculture and photovoltaic (PV) infrastructure and require careful design to balance crop performance and energy generation. A critical element of agrivoltaic design is the spatial and temporal distribution of irradiance and shade within and around PV arrays. To support research, planning, and stakeholder decision-making, we introduce the U.S. Agrivoltaic Shading Tool, a novel web-based application that delivers high-resolution irradiance and photosynthetically active radiation (PAR) modeling for ten standardized PV configurations across the conterminous United States. The tool leverages the National Laboratory of the Rockies (NLR) System Advisor Model (SAM) to perform detailed irradiance simulations, using meteorological data from the National Solar Radiation Database (NSRDB). Outputs include seasonal, monthly, weekly, and diurnal patterns of available sunlight, amount of shade, irradiance, and PAR at ground level within agrivoltaic system footprints. For a user's selected location, these results are visualized through interactive visualizations, heatmaps, and time-series plots, designed to be accessible to both technical and non-technical users. In addition to facilitating rapid spatial exploration of agrivoltaic light environments, the tool will offer seamless integration with the InSPIRE Agrivoltaics Design and Analysis Model (ADAM). This optional workflow will allow users to port selected site and configuration parameters into a more advanced modeling environment for further customization of structural layouts, crop-system compatibility, power generation, and technoeconomic performance. Finally, to promote open science, the entire dataset will be hosted and available for open access through the OpenEI platform. By standardizing and disseminating high-quality irradiance data and design tools, the U.S. Agrivoltaic Shading Tool supports a wide range of users, including researchers, landowners, energy developers, and policymakers, in evaluating the agronomic and energetic feasibility of agrivoltaic systems across the United States.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Machine learning pipeline for denoising low signal-to-noise ratio and out-of-distribution transmission electron microscopy datasets

High-resolution transmission electron microscopy (HRTEM) is crucial for observing material’s structural and morphological evolution at Angstrom scales, but the electron beam can alter these processes. Devices such as CMOS-based direct-electron detectors operating in electron-counting mode can be utilized to substantially reduce the electron dosage. However, the resulting images often lead to a low signal-to-noise ratio, which requires frame integration that sacrifices temporal resolution. Several machine learning (ML) models have been recently developed to successfully denoise HRTEM images. Yet, these models are often computationally expensive, and their inference speeds on GPUs are outpaced by the imaging speed of advanced detectors, precluding in situ analysis. Furthermore, the performance of these denoising models on datasets with imaging conditions that deviate from the training datasets has not been evaluated. To mitigate these gaps, we propose a new self-supervised ML denoising pipeline specifically designed for time-series HRTEM images. This pipeline integrates a blind-spot convolution neural network with pre-processing and post-processing steps, including drift correction and low-pass filtering. Results demonstrate that our model outperforms various other ML and non-ML denoising methods in noise reduction and contrast enhancement, leading to improved visual clarity of atomic features. Additionally, the model is drastically faster than U-Net-based ML models and demonstrates excellent out-of-distribution generalization. The model’s computational inference speed is in the order of milliseconds per image, rendering it suitable for application in in-situ HRTEM experiments.

36 MATERIALS SCIENCE↗

The South Pole Telescope AGN Monitoring Campaign: First Release of SPTpol Bright AGN Light Curves

The South Pole Telescope (SPT) collaboration has recently embarked upon a campaign to monitor the brightness of a sample of active galactic nuclei (AGN), both in real time and in archival SPT data. The original design of the SPT was optimized for observations of the cosmic microwave background (CMB) at arc-minute and larger angular scales, and it has been used for this purpose for nearly twenty years, using three generations of CMB cameras. Recently it has been recognized that data from CMB experiments have the potential to be used for AGN monitoring. In this paper, we present the first public release of data from a full sample of SPT-monitored AGN, comprising 158 AGN light curves and associated data from the SPTpol camera, which was operational from 2012-2016. These light curves were created using observations from the SPTpol 500 deg$^{2}$ survey, in which the instrument was used to scan a 500 deg$^2$ patch of the sky several times per day with detectors sensitive to radiation in bands centered at 90 and 150 GHz. We provide a comprehensive description of the observations, the data processing methods, and the resulting light curve catalog. As an example of analyses that these data enable, we searched for a correlation between variability and spectral index, and we looked for ``bluer-when-brighter'' trends in the sample. Our analysis finds $> 10 σ$ correlation between fractional intrinsic variance and mean spectral index in the sample, but no significant evidence for bluer-when-brighter trends. The datasets from this study can be accessed through the SPT Treasury Record of AGN With Historical Activity and Time-Series or STRAWHAT catalog. This initial data release includes SPTpol light curves at 90 and 150 GHz, focusing on total intensity. In later updates, SPTpol polarization data and new observations from the SPT-3G instrument at 90, 150, and 220 GHz will be included.

Hood, J.C., II [Chicago U., KICP; Chicago U., Astr↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database↗

Integrating very-high-resolution imagery, Sentinel-2 time-series data, and machine learning to map shrub fractional abundance across arid and semi-arid ecosystems in China

Shrub fractional abundance (SFA), the proportion of shrub cover per unit area, serves as a critical indicator of environmental aridity and ecosystem health in arid and semi-arid regions, particularly across the Mongolian steppe. However, large-scale SFA mapping in Mongolian steppe ecosystems remains challenging due to the small crown size of shrubs, their sparse distribution, and spectral overlap with coexisting low vegetation (e.g., grasses and herbs), which hinders accurate detection using coarser-resolution satellite data or traditional field surveys. To address these challenges, we developed a two-step approach that integrates very-high-resolution (VHR) imagery, time-series Sentinel-2 data, and deep learning techniques. First, we generated high-accuracy benchmark maps of individual shrub crowns from 0.5 m VHR imagery by combining manual segmentation with a hybrid deep learning framework (Dino V2 and convolutional neural networks). Second, we used these shrub crown maps as training data to build an XGBoost model for predicting SFA from 20 m Sentinel-2 time-series data, leveraging phenological information to improve estimation. We validated our approach across 70 sites (1km 2 each) in the Inner Mongolia Autonomous Region, which is representative of Mongolian steppe ecosystems. From VHR imagery, we mapped 1.31 million shrub crowns with an accuracy of R 2 = 0.92. Scaling up with Sentinel-2 data yielded regional SFA maps with an R 2 = 0.60. Further SHAP (SHapley Additive exPlanations) analysis on the developed XGBoost model revealed that phenological metrics (particularly observations in early-May, mid-July, and late-September), which distinguish shrub phenology from that of other land cover types (e.g., grasses and bare soil), were the most influential predictors of SFA. Finally, our regional SFA maps uncovered unimodal relationships between shrub distribution and climate variables, peaking at mean annual minimum temperatures near 0 °C and annual precipitation around 200 mm. Collectively, these findings demonstrate how the integration of multi-source remote sensing and machine learning can overcome historical limitations in SFA mapping, enabling accurate, spatially continuous assessments across vast Inner-Mongolian steppe ecosystems. Our framework has the potential to be applied to other steppe ecosystems and dryland ecosystems across the Mongolian steppe and beyond, offering a foundation for improved monitoring and ecological impact assessments in the face of global climate changes.

Arid and semi-arid landscapes↗

Knowledge Graph for End-to-End Traceability of an Integrated Human-Earth System Model

Integrated human-Earth system models inform energy-water-land system dynamics and policies, yet their results are difficult to trace through input-data, model structure, scenario configurations, and solved outputs. Because this information is siloed across disconnected artifacts, process-based IAMs have historically lacked a unified, queryable representation. Such lack of traceability prevents researchers from systematically isolating the multi-sector drivers of complex outcomes (such as tracing water-scarcity results back to distant energy-system dynamics) or conducting holistic uncertainty attribution across hundreds of interacting parameters. To address this concern, our work documents the software engineering process of a knowledge graph that unifies these four layers for the Global Change Analysis Model (GCAM-USA_Reference scenario, GCAM v9.1). The graph was built as a relational property graph in DuckDB from the run’s own artifacts: the input-preparation dependency map (gcamdata chunk map), the model’s XML input files, the run configuration, and the results database (BaseX), successfully mapping the model’s declared structure. The resulting graph comprises 204,321 nodes and 1,687,814 edges across 16 node types and 15 edge types, with approximately 16.3 million time-series values stored separately to maintain structural efficiency. To ensure representation fidelity, every edge carries an epistemic-status annotation recording the warrant for the relationship (structural, provenance, dependency, or model-derived), and a machine-readable provenance ledger classifying the origin of every schema element. Evaluation against a fixed five-benchmark suite with locked baselines reports zero structural orphans, zero dangling edge endpoints, and 100% of output-producing technologies traceable to raw input files. Two interactive interfaces present the graph, including a serverless browser application built on DuckDB-Wasm. By establishing the first end-to-end provenance framework for an IAM, this work enables researchers and scientists to systematically audit complex policy scenarios, debug model structures, and trace policy-relevant outputs to their data origins in real time.

Artifical Intelligence↗