Search NASA⌕ Search

SEARCH · Search NASA

Results for “hierarchical data change analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Assessing data change in scientific datasets

Summary Scientific datasets are growing rapidly and becoming critical to next‐generation scientific discoveries. The validity of scientific results relies on the quality of data used and data are often subject to change, for example, due to observation additions, quality assessments, or processing software updates. The effects of data change are not well understood and difficult to predict. Datasets are often repeatedly updated and recomputing derived data products quickly becomes time consuming and resource intensive and may in some cases not even be necessary, thus delaying scientific advance. Despite its importance, there is a lack of systematic approaches for best comparing data versions to quantify the changes, and ad‐hoc or manual processes are commonly used. In this article, we propose a novel hierarchical approach for analyzing data changes, including real‐time (online) and offline analyses. We employ a variety of fast‐to‐compute numerical analyses, graphical data change representations, and more resource‐intensive recomputations of a subset of the data product. We illustrate the application of our approach using three scientific diverse use cases, namely, satellite, cosmological, and x‐ray data. The results show that a variety of data change metrics should be employed to enable a comprehensive representation and qualitative evaluation of data changes.

97 MATHEMATICS AND COMPUTING↗

Monitoring Change Through Hierarchical Segmentation of Remotely Sensed Image Data

NASA's Goddard Space Flight Center has developed a fast and effective method for generating image segmentation hierarchies. These segmentation hierarchies organize image data in a manner that makes their information content more accessible for analysis. Image segmentation enables analysis through the examination of image regions rather than individual image pixels. In addition, the segmentation hierarchy provides additional analysis clues through the tracing of the behavior of image region characteristics at several levels of segmentation detail. The potential for extracting the information content from imagery data based on segmentation hierarchies has not been fully explored for the benefit of the Earth and space science communities. This paper explores the potential of exploiting these segmentation hierarchies for the analysis of multi-date data sets, and for the particular application of change monitoring.

Tilton, James C.↗

A hierarchical evaluation framework for assessing climate simulations relevant to the energy-water-land nexus (Final Report)

The overarching goal was to construct a hierarchy of new and well-tested metrics and analysis tools that support both fundamental and use-inspired research. Motivating this goal was a convergence of needs from climate scientists and stakeholders alike for a systematic, robust framework of model evaluation and diagnosis to provide scientific insights, inform model development, support best practices for the use of climate model outputs, and facilitate communication of climate information in the evolving landscapes of multi-model, multi-resolution, and large ensemble simulations that generate terabytes of data for any single climate run. Steps toward attaining these goals benefited from expertise and established capability of the project team, which included leadership of the North American Regional Climate Change Assessment Program (NARCCAP) and the Coordinated Regional Downscaling Experiment (CORDEX), development of hierarchical model evaluation approaches, and successful research in the analysis and diagnosis of climate model skill, as well as the understanding and modeling of regional climate processes in North America. As part of the overarching goal, the project worked to disseminate a suite of methodologies, algorithms, and software components that the wider community can employ to advance climate science and applications. With rigorous demonstration, the evaluation framework and the mix of standard and high risk / high reward approaches helped form the basis for future development of a computationally enabled user-friendly system for community use.

17 WIND ENERGY↗

Statistical framework to assess long-term spatio-temporal climate changes: East River mountainous watershed case study

Abstract Evaluation of long-term temporal and spatial climatic change in mountainous regions is a critical challenge because of the interactive effects of multiple land and climatic factors and processes. Here we present the application of the statistical framework to the assessment of changes of climatic conditions, using data from 17 meteorological stations across the East River watershed near Crested Butte, Colorado, USA, and spanning the period from 1966 to 2021. The framework is developed based on (1) a time-series analysis of daily, monthly, and yearly averaged meteorological parameters (temperature, relative humidity, precipitation, wind speed, etc.), (2) evaluation and time series analysis of potential evapotranspiration (ET o ), actual evapotranspiration (ET), aridity index (AI), standard precipitation index (SPI) and standard precipitation-evapotranspiration index (SPEI), and (3) a temporal-spatial climatic zonation of the studied area based on the hierarchical clustering and PCA analysis of the SPEI, because the SPEI can be considered an integrative characteristic of the changes of climatic conditions. The Budyko model, with the application of the Penman–Monteith equation for the estimation of ET o , was used to determine the ET. The time series analysis of the AI is used to identify the periods with energy limited and water limited conditions. Hierarchical clustering of site locations for the three temporal segments of the SPEI showed a significant temporal-spatial shifts, indicating that dynamic climatic processes drive zonation patterns. Therefore, the watershed climatic zonation requires periodic re-evaluation based on the structural time series analysis of meteorological and water balance data.

54 ENVIRONMENTAL SCIENCES↗

Ion Clusters Reveal the Sources, Impacts, and Drivers of Freshwater Salinization

Population growth, land use change, climate change, and natural resource extraction are driving the salinization of freshwater resources worldwide. Reversing these trends will require data-centric approaches that identify salt sources, environmental drivers, and ecosystem responses. In this study, we applied principal component analysis and hierarchical clustering to identify ion covariance patterns, or “ion clusters,” in Broad Run, an urban stream in the Mid-Atlantic United States. These clusters correspond to distinct hydrologic regimes and reveal specific salinization risks: (1) phosphorus pollution mobilized during summer storms (Cluster 1); (2) elevated concentrations of sulfate and bicarbonate during baseflow (Cluster 2), likely reflecting groundwater discharge; and (3) elevated specific conductance and sodium, chloride, and potassium ion concentrations during snowmelt and rain-on-snow events (Cluster 3), driven by deicer and anti-icer wash-off. These ion fingerprints offer a transferable framework for diagnosing salt sources, assessing ecological risk, and identifying management targets. Our findings underscore the need for next-generation stormwater infrastructure and smart growth policies to protect aquatic life in rapidly urbanizing watersheds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Utilizing Hierarchical Segmentation to Generate Water and Snow Masks to Facilitate Monitoring Change with Remotely Sensed Image Data

The hierarchical segmentation (HSEG) algorithm is a hybrid of hierarchical step-wise optimization and constrained spectral clustering that produces a hierarchical set of image segmentations. This segmentation hierarchy organizes image data in a manner that makes the image's information content more accessible for analysis by enabling region-based analysis. This paper discusses data analysis with HSEG and describes several measures of region characteristics that may be useful analyzing segmentation hierarchies for various applications. Segmentation hierarchy analysis for generating landwater and snow/ice masks from MODIS (Moderate Resolution Imaging Spectroradiometer) data was demonstrated and compared with the corresponding MODIS standard products. The masks based on HSEG segmentation hierarchies compare very favorably to the MODIS standard products. Further, the HSEG based landwater mask was specifically tailored to the MODIS data and the HSEG snow/ice mask did not require the setting of a critical threshold as required in the production of the corresponding MODIS standard product.

Tilton, James C.↗

Traffic safety analysis and model updating for freeways using Bayesian method

Freeway crash prediction models are the basic of traffic safety research, yet crash occurrence and the influencing factors change over time. In order to make sure the implemented safety models fit the current traffic environment, this study conducts a comparative analysis of 2017 and 2020 datasets collected from freeways in Suzhou, China. Herein, considering the spatial correlation among analysis units and the hierarchical data structure, a Bayesian conditional autoregressive negative binomial (CAR-NB) model and a Bayesian hierarchical CAR-NB (HCAR-NB) model were used to explore the safety influencing factors, and a traditional NB model was developed for further comparison. To update the HCAR-NB model from 2017 to 2020, Bayesian inference with informative priors was used to improve its goodness of fit and efficiency. Preliminary results showed that 1) the HCAR-NB model outperformed the NB model and CAR-NB model in prediction accuracy, and 2) the number of crashes was significantly correlated with average speed, speed variance, road segment length, number of lanes, and presence of ramps. The potential for safety improvement (PSI) method was applied to the modeling results to identify hotspots for the two years. The results confirmed that the hotspots spatiotemporally shifted among the freeways. The proposed crash prediction model and updating method are expected to assist implementation of informed countermeasures for freeway safety improvement.

97 MATHEMATICS AND COMPUTING↗

The Mars 2020 Ground Data System Architecture

The Mars 2020 Mission’s primary objective is to collect 20 geographically unique samples during its prime mission of one and a quarter Martian years, or just over 2 Earth years. Mission planners determined the project needed to develop a system that would enable the operations team to analyze engineering and science data, make science decisions, select viable rover targets at a millimeter resolution and validate an uplink bundle for a car sized rover with more complex science instruments than any previous Mars surface mission. All this had to be done within a five hour time frame. Doing this with a small team would be a challenge, but this had to be accomplished by a large team of engineers and scientists located across North America and Europe. Achieving this level of operational efficiency was unheard of in the prime mission. In addition, the mission had another set of requirements that had nothing to do with surface operations; the Mars 2020 Ground Data System (GDS) was also expected to comply with a new set of security requirements to keep up with the ever changing cybersecurity landscape. The Mars 2020 Ground Data System (GDS) is a re-architected version of the Mars Science Laboratory GDS. The primary goal was to integrate the lessons learned from previous Mars surface missions, accommodate a set of new requirements and capabilities required to ensure mission success, and comply with a new set of cybersecurity controls. The new architecture includes several unique qualities including a data lake, language-agnostic system-wide event-based operations, containerization, automated deployment, network segmentation, infrastructure-as-code, API-driven interfaces, and the first Mars surface GDS to operate primarily in the cloud. The new architecture enabled greater access to the system’s data, tighter integration with the operations team, and a higher level of traceability. The availability of the data also enabled a new set of capabilities previously not possible on surface missions. These new capabilities include an autonomous data to information, pipeline for downlink analysis, horizontal scaling of science data processing capabilities, autonomous round trip data tracking of science and engineering data, integration of flight system state into the tactical planning cycle, high fidelity targeting utilizing kinematic data, and hierarchical image and 3d meshes data representations. This paper will introduce the requirements for the Mars 2020 Mission, the heritage architecture, and the rationale for the changes to achieve the new architecture. The paper will continue to describe the fundamental changes made to the GDS architecture, how these changes enabled a more tightly integrated GDS, and the new capabilities that were enabled by the new architecture. The paper will conclude with the lessons learned from the process of rearchitecting a heritage GDS system and from the first 200 days of operations supporting over 800 users from around the world.

Lopez-Roig, Reynaldo↗

Spotlight: efficient automated global optimization in rietveld analysis of diffraction data

Performing reliable Rietveld analysis on tens or hundreds of powder diffraction datasets from parametric or time-resolved experiments often poses a bottleneck in extracting meaningful results from the data. While automated analysis of data has recently been demonstrated, high temperature annealing studies, during which phase transformations occur and lattice parameters may change due to repartitioning of elements, are prime examples where automation by a simple phase identification from a database of room temperature structures or automation by sequential refinements is likely to fail. To enable reliable, efficient, automated Rietveld analysis, we present a Python package named Spotlight , building on established Rietveld packages such as MAUD, GSAS , or GSAS-II , which extends the refinement of best fit parameters to a global optimization using an ensemble of optimizers leveraging hierarchical parallel execution on high-performance computing clusters. Spotlight further enables the efficient design of refinement plans through the iterative automated machine-learning of a surrogate for the refinement on which the global optimizations are performed until results from the surrogate converge to the response surface data. We demonstrate Spotlight with the analysis of uranium molybdenum and Ti–6Al–4V datasets, as well as in two open-source tutorials analyzing aluminium oxide and lead sulphate.

36 MATERIALS SCIENCE↗

PaleoSTeHM v1.0: a modern, scalable spatiotemporal hierarchical modeling framework for paleo-environmental data

Abstract. Geological records of past environmental change provide crucial insights into long-term climate variability, trends, non-stationarity, and nonlinear feedback mechanisms. However, reconstructing spatiotemporal fields from these records is statistically challenging due to their sparse, indirect, and noisy nature. Here, we present PaleoSTeHM, a scalable and modern framework for spatiotemporal hierarchical modeling of paleo-environmental data. This framework enables the implementation of flexible statistical models that rigorously quantify spatial and temporal variability from geological data while clearly distinguishing measurement and inferential uncertainty from process variability. We illustrate its application by reconstructing temporal and spatiotemporal paleo-sea-level changes across multiple locations. Using various modeling and analysis choices, PaleoSTeHM demonstrates the impact of different methods on inference results and computational efficiency. Our results highlight the critical role of model selection in addressing specific paleo-environmental questions, showcasing the PaleoSTeHM framework's potential to enhance the robustness and transparency of paleo-environmental reconstructions.

58 GEOSCIENCES↗

Topological Data Analysis for Particulate Gels

Soft gels, formed via the self-assembly of particulate materials, exhibit intricate multiscale structures that provide them with flexibility and resilience when subjected to external stresses. Here, this work combines particle simulations and topological data analysis (TDA) to characterize the complex multiscale structure of soft gels. Our TDA analysis focuses on the use of the Euler characteristic, which is an interpretable and computationally scalable topological descriptor that is combined with filtration operations to obtain information on the geometric (local) and topological (global) structure of soft gels. We reduce the topological information obtained with TDA using principal component analysis (PCA) and show that this provides an informative low-dimensional representation of the gel structure. We use the proposed computational framework to investigate the influence of gel preparation (e.g., quench rate, volume fraction) on soft gel structure and to explore dynamic deformations that emerge under oscillatory shear in various response regimes (linear, nonlinear, and flow). Our analysis provides evidence of the existence of hierarchical structures in soft gels, which are not easily identifiable otherwise. Moreover, our analysis reveals direct correlations between topological changes of the gel structure under deformation and mechanical phenomena distinctive of gel materials, such as stiffening and yielding. In summary, we show that TDA facilitates the mathematical representation, quantification, and analysis of soft gel structures, extending traditional network analysis methods to capture both local and global organization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Robust Carbon Dioxide Plume Imaging Using Joint Tomographic Inversion of Seismic Onset Time and Distributed Pressure and Temperature Measurements (Final Report)

We develop and demonstrate rapid and cost-effective methodologies for spatiotemporal tracking of CO2 plumes during geologic sequestration using joint inversion of seismic data and distributed pressure and temperature measurements. Key elements of our methodology are: (a) a computationally efficient approach to pressure and temperature propagation, (b) analysis of time lapse seismic data using a novel ‘seismic onset time’ approach to detect fluid front propagation, and (c) data assimilation and uncertainty assessment via joint inversion of pressure, temperature and time lapse seismic data, and (d) validating the numerical tomographic inversion using a CO2 injection demonstration projects, specifically data collected from the from the Petra Nova Parish Holdings CCUS project in the West Ranch Field, Texas and the Chester-16 reef CO2 injection site in Northern Michigan which is part of the DOE Midwestern Carbon Sequestration Project. The research team is led by Texas A&M University and includes Battelle as a subcontractor with support from Shell, Anadarko, Chevron and JX Nippon. A carbon dioxide (CO2) water-alternating-gas (WAG) pilot was conducted to gain insights into tertiary oil recovery potential via CO2 flood in the West Ranch Field as part of the Petra Nova project, the world’s largest post-combustion CO2 capture and utilization initiative. With a fluvial formation geology and large contrasts in permeability, this is a challenging and novel application of CO2 enhanced oil recovery (EOR). We build a predictive dynamic model of the subsurface that incorporates the multiphase and compositional data acquired during the pilot operation. The calibrated model is used for the carbon dioxide plume imaging. The study began with an initialization of the pilot sector model extracted from a calibrated full-field model. The pilot model calibration follows a two-step hierarchical workflow. First, we performed a large-scale update of the permeability distribution by integrating available bottomhole pressure and multiphase production data. In the second step, local permeability field is fine-tuned using a streamline-based method to match CO2 breakthrough times at the producers. The predictive capability of the calibrated model was verified through two blind validation tests: (1) the model showed good agreement with saturation logs acquired at two observation wells; and (2) the model reproduced the CO2 recovery as a fraction of the injected CO2. The use of seismic onset times has shown great promise for integrating near-continuous seismic surveys for updating geologic models. In this study, we analyze the impact of seismic survey frequency on the onset time approach aiming to extend the application of onset time to infrequent seismic surveys. In addition, we quantitatively examine the nonlinearity of the onset time method and compare it to the commonly used amplitude inversion method. We carry out a sensitivity analysis of seismic survey frequency based on the complete seismic survey data (over 175 surveys) of steam injection in a heavy oil reservoir (Peace River Unit) in Canada. Our results show that an adequate onset time map can be obtained from the infrequent seismic surveys by interpolation between seismic surveys as long as there is no change in the dominant underlying physics between the successive surveys. The study also shows that nonlinearity of the onset time method can be -smaller than that of the amplitude inversion method by several orders of magnitude. Application to the Brugge benchmark case shows that the onset time method obtains comparable permeability update as the traditional seismic amplitude inversion method with faster computation and improved convergence characteristics. We extend the streamline-based data integration approach to incorporate distributed temperature sensor (DTS) data using the concept of thermal tracer travel time. Then, a hierarchical workflow composed of evolutionary and streamline methods is employed to jointly history match the DTS and pressure data. Finally, CO2 saturation and streamline maps are used to visualize the CO2 plume movement during the sequestration process. The hierarchical workflow is applied to a carbon sequestration project in a carbonate reef reservoir within the Northern Niagaran Pinnacle Reef Trend in Michigan, USA. The monitoring data set consists of distributed temperature sensing (DTS) data acquired at the injection well and a monitoring well, flowing bottom-hole pressure data at the injection well, and time-lapse pressure measurements at several locations along the monitoring well. The history matching results indicate that the CO2 movement is mostly restricted to the intended zones of injection which is consistent with an independent warm-back analysis of the temperature data. In addition to employing simulation models and inverse methods for CO2 plume imaging, we also initialized a data-driven technology for detecting inter-well connectivity based on production and pressure data. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO2 EOR projects utilizing the water-alternating-gas (WAG) process. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. Texas A&M University, the lead organization in the project, was primarily responsible for the development of tomographic approaches for CO2 plume mapping in conjunction with distributed pressure, temperature and seismic onset time data. Battelle, as a subcontractor, was primarily responsible for the development of analytical and empirical methods for analyzing transient injection rate and pressure data from point/line sources such as injection and monitoring wells. An additional area of emphasis for Battelle was the use of machine learning for such tasks as inferring reservoir connectivity information from injection-production data, and identifying variable importance for machine learning-based proxy models developed from full-physics simulations. The two organizations also collaborated on the application of the tomographic inversion methodology for a field data set.

02 PETROLEUM↗

Keywords for All People: How Keyword Governance and Coordination with the NASA ESDIS Standards Coordination Office (ESCO) Improves GCMD Keywords for Discuvery and Use

The Global Change Master Directory (GCMD) Keywords, initiated over twenty years ago, are a hierarchical set of controlled Earth Science vocabularies that help ensure Earth science data, services, and variables are described in a consistent and comprehensive manner and allow for the precise searching of metadata and subsequent retrieval of data, services, and variables. GCMD keywords are periodically analyzed for relevancy and will continue to be refined and expanded in response to user needs. The periodic analysis is a result of successful coordination with the ESDIS Standards Coordination Office (ESCO), which is responsible for standards activities across ESDIS, and assists with providing valuable stakeholder and subject matter expert (SME) feedback on GCMD vocabularies. The ESCO is also turning to the GCMD to discuss how the keyword review process can be improved and more streamlined. In addition to the ESCO, keyword requests and feedback are also received through the GCMD Keyword Forum.

Tyler Stevens↗

Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem

Coastal terrestrial-aquatic interfaces (TAIs) are crucial contributors to global biogeochemical cycles and carbon exchange. A systematic evaluation of the interaction between coastal catchment properties and carbon dioxide (CO2) emission by soil respiration is significant for assessing carbon dynamics and predicting the future trajectory of atmospheric CO2 concentrations in coastal TAIs. The soil CO2 efflux in these transition zones is however poorly understood due to the high spatiotemporal dynamics of TAIs, as various sub-ecosystems in this region are compressed and expanded by complex influences of tides, changes in river levels, climate, and land use. We focus on the Chesapeake Bay region to (i) investigate the spatial heterogeneity of the coastal ecosystem and identify spatial zones with similar environmental characteristics based on the spatial data layers, including vegetation index (kNDVI), climate, landcover, diversity, topography, soil property, and relative tidal elevation; (ii) understand the primary driving factors affecting soil respiration within sub-ecosystems of the coastal ecosystem. Specifically, we employed hierarchical clustering analysis to identify spatial regions with distinct environmental characteristics, followed by the determination of main driving factors using Random Forest regression and SHapley Additive exPlanations. Maximum and minimum temperature are the main drivers common to all sub-ecosystems, while each region also has additional unique major drivers that differentiate them from one another. Precipitation exerts an influence on vegetated lands, while soil pH value holds importance specifically in forested lands. In croplands characterized by high clay content and low sand content, the significant role is attributed to bulk density. Wetlands demonstrate the importance of both elevation and sand content, with clay content being more relevant in non-inundated wetlands than in inundated wetlands. The topographic wetness index significantly contributes to the mixed vegetation areas, including shrub, grass, pasture, and forest. Additionally, our research reveals that dense vegetation land covers and urban/developed areas exhibit distinct soil property drivers. Overall, there is no one-size-fits-all approach to modeling carbon fluxes in coastal TAIs, and our study highlights the importance of further research and monitoring practices to improve our understanding of carbon dynamics and promote the sustainable management of coastal TAIs.

54 ENVIRONMENTAL SCIENCES↗

Central Atlantic regional ecological test site: A prototype regional environmental information system

The author has identified the following significant results. A comparison of photomorphic regions from an uncontrolled ERTS-1 mosaic of CARETS to land use areas on a map published in the National Atlas revealed close correlations in non-urban regions. Such regional scale analysis of ERTS-1 data has the potential for providing an economical sampling strategy for selecting sites for more detailed field measurements if other environmental variables can be correlated with patterns on ERTS-1 imagery. ERTS-1 imagery has also revealed for the first time the appearance of CARETS during the winter months. Investigators have identified extensive areas of conifers, which have previously been indistinguishable from deciduous vegetation. Imagery has also shown very clearly the extent of snow cover at a particular time over the region. The evaluation of ERTS-1 imagery used for the land use mapping of the shore zone of CARETS, has shown that the presence or absence of elements of an hierarchal system of shoreline landforms can help identify areas of potential rapid change. Changes in land use class distributions on the Barrier Islands signify the environmental response to natural and man-caused processes. Both environmental vulnerability and sensitivity can be estimated from the repetitive ERTS-1 coverage of long reaches of the CARETS coast. Results indicate potential applications to land use planning, management, and regional environmental quality analysis.

Alexander, R. H.↗

CONGO²: Scalable Online Anomaly Detection and Localization in Power Electronics Networks

Rapid and accurate detection and localization of electronic disturbances simultaneously are important for preventing its potential damages and determining potential remedies. Existing anomaly detection methods are severely limited by the low accuracy, the expensive computational cost and the need for highly trained personnel. There is an urgent need for a scalable online algorithm for in-field analysis of large-scale power electronics networks. Here in this paper, we propose a fast and accurate algorithm for anomaly detection and localization of power electronics networks: stratified colored-node graph (CONGO2). This algorithm hierarchically models the change of correlated waveforms and then correlated sensors using the colored-node graph. By aggregating the change of each sensor with its neighbors’ inputs, we can spontaneously identify and localize the anomaly that cannot be detected by data collected from a single sensor. As our proposed method only focuses on the changes within a short time frame, it is highly computational efficient and only needs small data storage. Thus, our method is ideal for online and reliable anomaly detection and localization of large-scale power electronic networks. Compared to existing anomaly detection methods, our method is entirely data-driven without training data, highly accurate and reliable for wide-spectrum anomalies detection, and more importantly, capable of both detection and localization. Thus, it is ideal for infield deployment for large-scale power electronic networks. As illustrated by a distributed energy resources (DERs) power grid with 37-node, our method can effectively detect and localize various cyber and physical attacks.

42 ENGINEERING↗

Local-scale Arctic tundra heterogeneity affects regional-scale carbon dynamics

In northern Alaska nearly 65% of the terrestrial surface is composed of polygonal ground, where geomorphic tundra landforms disproportionately influence carbon and nutrient cycling over fine spatial scales. Process-based biogeochemical models used for local to Pan-Arctic projections of ecological responses to climate change typically operate at coarse-scales (1km 2 –0.5°) at which fine-scale (<1km 2 ) tundra heterogeneity is often aggregated to the dominant land cover unit. Here, we evaluate the importance of tundra heterogeneity for representing soil carbon dynamics at fine to coarse spatial scales. We leveraged the legacy of data collected near Utqiagvik, Alaska between 1973 and 2016 for model initiation, parameterization, and validation. Simulation uncertainty increased with a reduced representation of tundra heterogeneity and coarsening of spatial scale. Hierarchical cluster analysis of an ensemble of 21 st -century simulations reveals that a minimum of two tundra landforms (dry and wet) and a maximum of 4km 2 spatial scale is necessary for minimizing uncertainties (<10%) in regional to Pan-Arctic modeling applications.

54 ENVIRONMENTAL SCIENCES↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗