Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Reasoning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

ϕ → 3 π and ϕ π 0 transition form factor from Khuri-Treiman equations

This work studies the ϕ → 3 π decay and the ϕ → π 0 γ * transition form factor, utilizing the Khuri-Treiman formalism to account for analyticity, crossing, and unitarity. Using once-subtracted dispersion relations, we perform a simultaneous fit to the ϕ → 3 π Dalitz plot distribution and the ϕ → π 0 γ * measurements from the KLOE collaboration, finding good agreement with these experimental data. These results reaffirm the applicability of the Khuri-Treiman approach in the analysis of three-body decays. An interesting result is that the subtraction constant appearing in the equations is similar to a sum rule expectation, in contrast to analogous studies of ω → 3 π decays and ω → π 0 γ * , which shows significant deviations. Our results also provide a reasonable description of the trend of the transition form factor data from in the e + e − → ϕ π 0 scattering region. These intriguing theoretical differences between the decays of ϕ and ω could encourage further experimental measurements to assess the discrepancies and refine the theoretical predictions.

Lorenzo, A. G. [Instituto de Física Corpuscular; U↗

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

Measurement of coherent exclusive J/ψ → μ+μ− production in ultraperipheral Pb+Pb collisions at sNN=5.36 TeV with the ATLAS detector

The ATLAS experiment has performed a measurement of coherent exclusive J/ψ → μ+μ− production in ultraperipheral Pb+Pb collisions at sNN=5.36$$ \sqrt{s_{\textrm{NN}}}=5.36 $$ TeV. The data was recorded at the Large Hadron Collider (LHC) during 2023, and corresponds to an integrated luminosity of 79 μb−1. Exclusive J/ψ candidates were selected with a dedicated track-sensitive trigger based on the ATLAS transition radiation tracker. The analysis involves reconstruction of the dimuon invariant mass based on muon tracks from the inner detector, as the muon transverse momentum range of interest precludes the use of the standard muon reconstruction and identification algorithms. Differential cross sections are measured as a function of J/ψ rapidity and are compared with theoretical predictions. After extrapolation to sNN=5.02$$ \sqrt{s_{\textrm{NN}}}=5.02 $$ TeV, they are also compared with previous measurements performed by other experiments using data from LHC Run 2. While the results agree reasonably well with theoretical predictions, they are in tension with previous Run-2 results for the central rapidity region.

Aad, G↗

Annual Summary Report (FY 2025) Performance Assessment for the Integrated Disposal Facility

The purpose of this Annual Summary Report (ASR) for fiscal year (FY) 2025 is to evaluate the continued adequacy of the Integrated Disposal Facility (IDF) Performance Assessment (PA) and Disposal Authorization Statement (DAS). This report consolidates relevant monitoring data, modeling analyses, and regulatory reviews to demonstrate a reasonable expectation that the PA objectives and performance measures will be met, as required under DOE O 435.1, Radioactive Waste Management. The ASR follows the guidance in DOE-STD-5002-2017, Disposal Authorization Statement and Tank Closure Documentation, which provides a framework for maintaining the validity of the DAS through periodic assessment of facility performance and compliance with waste disposal requirements.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

A Perspective on Traditional and Data Driven Electrochemical Modeling and Analysis

To understand the behavior of electrochemical systems, we need to reduce the dimensionality of the measured current-voltage-time (I-V-t) data by fitting models, thus enabling us to analyze and compare the governing physics. Traditionally, the process for this is an 'expert first' approach: defining the model and its explicit assumptions based on inductive reasoning or empirical observation, fitting small portions of the I-V-t data where assumptions are most valid or carefully designing experiments to enforce key assumptions, and then interpreting the model parameters. However, modern data-driven methods enable a new paradigm: a 'data first' approach, where the latent behaviors governing the system's measured response are identified directly using machine-learning models that optimize both model structure and parameters from the I-V-t data, guaranteeing that the learned model explains as much of the observed system response as possible. After model identification, the model can then be interrogated by an expert to connect observed behaviors with underlying physics. This talk will review several different types of electrochemical analysis (electrochemical impedance, differential voltage-capacity, electrochemical kinetics) and compare the traditional and data-driven methods for analyzing the data.

42 ENGINEERING↗

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE↗

Visualization for Insight and Data Analysis in Energy Research

This talk explores how advanced visualization technologies are transforming analytical reasoning and knowledge discovery in energy research, drawing on recent work at the National Laboratory of the Rockies' Computational Science Center. Through a series of scientific case studies, we demonstrate how immersive and high-resolution visualization environments enable scientists and engineers to identify previously unseen patterns and features - insights that often remain hidden in traditional desktop-based analysis. By embedding richer information into interactive analytics tools, these approaches support the exploration of complex, multivariate parameter spaces, where interaction itself catalyzes understanding. Beyond capability, we emphasize the critical role of visualization design grounded in perception and cognition, showing how visual encodings directly influence analytical outcomes. Spanning applications from materials science to integrated energy systems, these visualization approaches accelerate innovation and improve decision-making by enabling deeper, more reliable insight into increasingly complex energy data.

97 MATHEMATICS AND COMPUTING↗

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Persistent Current Simulation for CCT Testing Magnet Used in EIC

The Electron-Ion Collider (EIC), a powerful new facility to be built in the United States at the U.S. Department of Energy's Brookhaven National Laboratory in collaboration with Thomas Jefferson National Accelerator Facility, will explore the most fundamental building blocks of nearly all visible matter. Here, there are many different types of superconducting magnets near the interaction region (IR) of EIC. Due to space constraints and special lattice requirements, Tapered CCT (canted-cosine theta) magnets have been used for EIC. At beam injection, the magnetic field is only ~5.5% of the maximum operating field. Considerable field errors will be generated from persistent current in superconducting strands even using very fine filament for those superconductors. A tapered CCT demonstrator magnet has been built and tested successfully at BNL since July 2020 to evaluate the key technologies for future tapered CCT magnets. In October 2023, BNL team also measured the persistent current in this demonstrator magnet. To validate the persistent current simulation methods for CCT magnets in EIC, this paper used a full 3D Opera Model and measured magnetization data from superconducting strand for the simulation. Simulation results showed reasonable agreement with recent measurement results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

High-temperature seals for supercritical carbon-dioxide (sCO 2 ) turbines (Final Report)

This is the final report for project DE-FE0031924 titled “High-temperature seals for supercritical carbon-dioxide (sCO 2 ) turbines.” The report provides a summary of the entire project efforts from October 2020 through December 2024 including the high-temperature commercial dry gas seal (DGS) tests and thermal modeling of Task 2, as well as the high-temperature, large-diameter seal design and high-temperature tests of large-diameter seals in Task 3. A key outcome of Task 2 was the testing completion of specially instrumented commercial DGS in the GE-SwRI Apollo sCO 2 compressor (27,000 rpm). Test data from the DGS showed elevated temperatures upwards of 350 o F, which are close to the higher operating temperature limit of the DGS. The temperature measurements provide insight into the expected thermal loads on DGS operating in high-speed sCO 2 compressor and provided test data for validation of an in-house thermal model of the compressor/seal. Under Task 2.0, this report also presents the development of a steady-state conjugate heat-transfer model of the DGS operating in the sCO 2 compressor – a first of its kind model for modeling heat transfer of sCO 2 in an actual operating compressor. The findings of the thermal model show a reasonable match between temperature predictions of the model and the measured temperature data, also pointing out the validity of the approach and assumptions made in modeling the flows, heat transfer coefficients and windage modeling in the rig. Under Task 3.0, this report presents the preliminary design of a large-diameter hybrid face seal (14 inch and 26-inch diameter) for field testing in a land-based GE turbine. The preliminary seal design effort presented in this report under project DE-FE0031924 builds on the development and successful laboratory testing for such large diameter hybrid face seal under the prior DE-FE0024007 project. Key aspects of seal fluid analyses with CFD, mechanical design considerations and assembly considerations in a land-based turbine are presented. Finally, under Task 3.0, this report also presents the continued high-temperature testing of the 14-inch diameter hybrid face seal developed previously under the DE-FE0024007 program. Specifically, test data demonstrating successful non-contact seal operation and seal effective leakage of 0.001-inch with seal inlet temperatures above 700 o F are presented in this report. Successful hybrid seal operation in a laboratory environment for a large diameter (14-inch) seal at temperatures above 700 o F is a major technological milestone for this technology.

01 COAL, LIGNITE, AND PEAT↗

Measurements of Gamow-Teller transitions from 59 Co via the 59 Co ⁢(𝑡, 3 He +𝛾) charge-exchange reaction and its application to the stellar electron-capture rates

Electron-capture reactions on iron-group nuclei play a crucial role in the late stages of massive star evolution. Since stellar evolution simulations depend on accurate electron-capture rates—which are highly sensitive to the detailed Gamow-Teller (GT) strength distributions—reliable theoretical models are essential. However, experimental data on GT strength distributions are scarce. High-resolution measurements are therefore vital for benchmarking and improving these theoretical calculations. To provide high-resolution data on Gamow-Teller strength distributions of iron-group nuclei and to compare these results with theoretical calculations within this mass region. Differential cross sections for the 59 Co ⁢(𝑡, 3 He)⁢ 59 Fe charge-exchange reaction at 115 MeV/u were measured using the S800 spectrometer. Furthermore, to resolve individual levels that are not distinguishable in the S800 particle singles data, coincident 𝛾 rays from the 59 Fe residual nucleus were detected by using the Gamma-Ray Energy Tracking In-beam Nuclear Array 𝛾-ray tracking array. Here, the Gamow-Teller transition strength distribution from the ground state of 59 Co to 59 Fe was extracted up to an excitation energy of 10 MeV. Additionally, transition strengths for several low-lying states were determined from coincident 𝛾-ray measurements. Electron-capture rates calculated using the present data indicate that these low-lying states contribute significantly to the overall rates in relevant stellar environments. The experimental results show reasonable agreement with theoretical predictions based on both shell-model and projected shell-model calculations. High-resolution data on Gamow-Teller strength distributions—particularly for individual low-lying states—are essential for accurately determining electron-capture rates in iron-group nuclei. Coincident 𝛾-ray measurements provide a powerful tool for obtaining such detailed information. While the present work demonstrates that shell-model calculations successfully reproduce the experimental results, such comparisons are scarce and more experimental data are desirable.

59 ≤ A ≤ 89↗

What Are Ontologies and When Should They Be Used?

Data without description is at best unusable, and at worst, misused. If we do not understand the assumptions and meaning of our data, we are unable to confidently use it. Data today is largely described within a database’s schema, detailing structure and primitive datatypes as part of a relational model, but if we require assurance some data value can be correctly evaluated alongside others beyond the immediate systems in which they are defined, a more portable, richer semantics is needed. Ontologies define knowledge unambiguously across systems and establish the means to reason upon said knowledge using logical inference. They model neutral domains of information rather than data definitions from software or databases that would only serve to enrich a single system’s idiosyncrasies. In this paper, we take a casual stance to explore what ontologies are, how they are built, why they are useful, and when they should be used.

97 MATHEMATICS AND COMPUTING↗

Archi: Agentic Operations at the CMS Experiment

We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them. An instance of Archi has been deployed for the Computing Operations team of the CMS experiment at CERN's LHC since February 2026 as a support agent for technical operators, offering retrieval and analysis capabilities by combining documentation, historical data, and live monitoring systems. We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels. The system proves effective at operational tasks, resolving real-world queries posed by CMS operators. We also observe that locally-hosted, open-weight models perform competitively, enabling fully private management of sensitive data.

Lugato, Pietro [MIT; CERN]↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗

Understanding Aitken Mode Aerosol Variability over the Southern Ocean and Antarctica: Insights from Cloud Condensation Nuclei Data

Aitken mode aerosol particles play an important role influencing cloud properties and sustenance, acting as a reservoir of potential cloud condensation nuclei against precipitation scavenging. However, there is limited data on Aitken mode aerosols. In this study, we develop a method to estimate Aitken mode aerosol concentrations and size distribution using cloud condensation nuclei measurements (CCN) and κ-Köhler theory. The performance of this method is evaluated using scanning mobility particle sizer (SMPS) data from recent field campaigns to demonstrate its skills and applicability. The method reasonably estimates Aitken- and accumulation-mode aerosol concentrations, achieving correlations of 0.7–0.9 with only modest biases (mean fractional bias within ±23% for Aitken mode and ±34% for accumulation-mode). This method is further applied to measurements collected over the Southern Ocean and Antarctica in recent years from multiple platforms, including ground sites, aircraft, and ships, to derive Aitken and accumulation-mode aerosol concentrations. Using the derived data, we examine the seasonal cycle, latitudinal variations, and vertical distribution of aerosols. Aitken mode aerosol concentrations are elevated over the Southern Ocean and Antarctica during the austral summer similar to the accumulation mode. In the austral summer, the free troposphere has more Aitken mode aerosols and fewer accumulation mode aerosols than the boundary layer, and thus likely serves as an important source of cloud-forming aerosol while also diluting the accumulation mode.

Kang, Litai [University of Washington] (ORCID:0000↗

Towards an IPv6-only WLCG: More successes in reducing IPv4

The Worldwide Large Hadron Collider Computing Grid (WLCG) community’s deployment of dual-stack IPv6/IPv4 on its worldwide storage infrastructure has been very successful. Dual-stack is not, however, a viable longterm solution; the HEPiX IPv6 Working Group has focused on studying where and why IPv4 is still being used, and how to flip such traffic to IPv6. The agreed end goal is to turn IPv4 off and run IPv6-only over the wide-area network to simplify both operations and security management.This paper reports our work since the CHEP2023 conference. Firstly, we present our campaign to deploy IPv6 on CPU services and Worker Nodes, with a deadline of end of June 2024. Then, the WLCG Data Challenge (DC24) performed in February 2024 was an excellent opportunity to observe the percentage of data transfers carried by IPv6. We observed the predominance of IPv6 in data transfers during DC24 and were able to understand yet more reasons for the use of IPv4 and areas for remedial action.The paper ends with the working group’s plans for moving WLCG to “IPv6- only”. One aspect of this is the possible automated use of IPv6-only clients configured with a customer-side translator, or CLAT, together with a deployment of NAT64 using what is often known as “IPv6-Mostly”, enabling IPv6-only sites to connect to non-WLCG IPv4-only services.

Attebury, Garhan [U. Nebraska, Lincoln]↗

Integrating machine learning interatomic potentials with hybrid reverse Monte Carlo structure refinements in RMCProfile

Structure refinement with reverse Monte Carlo (RMC) is a powerful tool for interpreting experimental diffraction data. To ensure that the under-constrained RMC algorithm yields reasonable results, the hybrid RMC approach applies interatomic potentials to obtain solutions that are both physically sensible and in agreement with experiment. To expand the range of materials that can be studied with hybrid RMC, we have implemented a new interatomic potential constraint in RMCProfile that grants flexibility to apply potentials supported by the Large-scale Atomic/Molecular Massively Parallel Simulator ( LAMMPS ) molecular dynamics code. This includes machine learning interatomic potentials, which provide a pathway to applying hybrid RMC to materials without currently available interatomic potentials. To this end, we present a methodology to use RMC to train machine learning interatomic potentials for hybrid RMC applications.

Cuillier, Paul↗