Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data processing methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42

Heterogeneous Point Set Transformers for Segmentation of Multiple View Particle Detectors

NOvA is a long-baseline neutrino oscillation experiment that detects neutrino particles from the NuMI beam at Fermilab. Before data from this experiment can be used in analyses, raw hits in the detector must be matched to their source particles, and the type of each particle must be identified. This task has commonly been done using a mix of traditional clustering approaches and convolutional neural networks (CNNs). Due to the construction of the detector, the data is presented as two sparse 2D images: an XZ and a YZ view of the detector, rather than a 3D representation. We propose a point set neural network that operates on the sparse matrices with an operation that mixes information from both views. Our model uses less than 10% of the memory required using previous methods while achieving a 96.8% AUC score, a higher score than obtained when both views are processed independently (85.4%).

Robles, Edgar E. [UC, Irvine (main)]↗

Impact of T - and ρ -dependent decay rates and new (n, γ ) cross-sections on the s process in low-mass asymptotic giant branch stars

Aims. We study the impact of nuclear input related to weak-decay rates and neutron-capture reactions on predictions for the slow neutron-capture process (s process) in asymptotic giant branch (AGB) stars. We provide the first database of surface abundances and stellar yields of the isotopes heavier than iron from the Monash models. Methods. We ran nucleosynthesis calculations with the Monash post-processing code for seven stellar structure evolution models of low-mass AGB stars with three different sets of nuclear inputs. The reference set has constant decay rates and represents the set used in the previous Monash publications. The second set contains the temperature and density dependence of β decays and electron captures based on the default rates of nuclear NETwork GENerator (NETGEN). In the third set, we further update 92 neutron-capture rates based on re-evaluated experimental cross sections from the ASTrophysical Rate and rAw data Library. We compare and discuss the predictions of the sets relative to each other in terms of isotopic surface abundances and total stellar yields. We also compare the results to isotopic ratios measured in presolar stardust silicon carbide (SiC) grains from AGB stars. Results. The new sets of models result in a ∼66% solar s-process contribution to the p-nucleus 152 Gd, confirming that this isotope is predominantly made by the s process. The nuclear input updates result in predictions for the 80 Kr/ 82 Kr ratio in the He intershell and surface 64 Ni/ 58 Ni, 94 Mo/ 96 Mo, and 137 Ba/ 136 Ba ratios that are more consistent with the corresponding ratios measured in stardust; however, the new predicted 138 Ba/ 136 Ba ratios are higher than the typical values of the SiC grains. The W isotopic anomalies are in agreement with data from the analyses of other meteoritic inclusions. We confirm that the production of 176 Lu and 205 Pb is affected by too large uncertainties in their decay rates from NETGEN.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine learning inversion from scattering for mechanically driven polymers

A machine learning inversion method is developed for analyzing scattering functions of mechanically driven polymers and extracting the corresponding feature parameters, which include energy parameters and conformation variables. The polymer is modeled as a chain of fixed-length bonds constrained by bending energy, and it is subject to external forces such as stretching and shear. We generate a data set consisting of random combinations of energy parameters, including bending modulus, stretching and shear force, along with Monte Carlo-calculated scattering functions and conformation variables such as end-to-end distance, radius of gyration and off-diagonal component of the gyration tensor. The effects of the energy parameters on the polymer are captured by the scattering function, and principal component analysis ensures the feasibility of the machine learning inversion. Finally, we train a Gaussian process regressor using part of the data set as a training set and validate the trained regressor for inversion using the rest of the data. The regressor successfully extracts the feature parameters.

Gaussian process regressors↗

Oxygen Storage Incorporated Into Net Power and the Allam–Fetvedt Oxy-Fuel sCO2 Power Cycle—Techno-Economic Analysis

Abstract With the planned future reliance on variable renewable energy, the ability to store energy for prolonged time periods will be required to reduce the disruption of market fluctuations. This paper presents a method to analyze a hybrid liquid-oxygen (LOx) storage/direct-fired supercritical carbon dioxide (sCO2) power cycle and optimize the economic performance over a diverse range of scenarios. The system utilizes a modified version of the NET Power process to produce energy when energy demand exceeds the supply while displacing much of the cost of the air separation unit (ASU) energy requirements through cryogenic storage of oxygen. The model uses marginal cost of energy data to determine the optimal times to charge and discharge the system over a given scenario. The model then applies ramp rates and other time-dependent factors to generate an economic model for the system without storage considerations. The size of the storage system is then applied to create a realistic model of the plant operation. From the real plant operation model, the amount of energy charged and discharged, the capital expenditures (CAPEX) of each system, energy costs and revenue and other parameters can be calculated. The economic parameters are then combined to calculate the net present value (NPV) of the system for the given scenario. The model was then run through the SMPSO genetic algorithm in Python for a variety of geographic regions and large-scale scenarios (high solar penetration) to maximize the NPV based on multiple parameters for each subsystem. The LOx storage requirements will also be discussed.

Engineering↗

Common Electric Power Transmission System Model JSON Schema Specification

The Common Electric Power Transmission System Model (CTM) is an intuitive, extensible, language-agnostic, and error-resistant specification of electric power network components parameter names and units, and relation between components, intended for use by the research community developing new computational methods for power systems operations and simulation. Power system datasets following the CTM specification can be read as dictionaries and manipulated in that form in most programming languages (e.g., Python, Julia, C++). This standard data structure in CTM makes it easy to work in multiple power systems domains (e.g., economic operation, reliability assessment, electricity markets, stability assessment, etc.) without requiring conversions between use-case-specific file formats with information loss in the process. This repository specifies CTM as a JSON Schema, provides documentation, derivate (code-generated) implementations of CTM, and example data and usage of the schema for important use cases.

Aravena Solis, Ignacio↗

Qualification of Digitized Legacy Fast Reactor Data

The Integral Fast Reactor (IFR) fuel compatibility test program (1984-1994) included a variety of fuel pin examinations conducted at the Hot Fuel Examination Facility (HFEF) and the Alpha-Gamma Hot Cell Facility (AGHCF). Hard copy data records of these examinations have been recovered, scanned, and preserved in PDF format. Many hard copy records are now qualified in accordance with an NRC-approved Quality Assurance Program Plan (QAPP), and there is an ongoing effort to qualify additional legacy records. This legacy fuel performance data is vital to support design and licensing of fast reactors with validation of state-of-the-art codes and advanced methods for design and analysis. Stakeholders can most easily utilize this data when the PDF scans have been converted into digital data tables. However, qualification of the scanned hard copy data does not qualify the digital data file resulting from the digitization of the data contained in the record; the subject matter expert (SME) must make a review of the digitized data table as well before it can be designated as qualified. This report outlines a peer review process to qualify the digital data file(s), typically in CSV format, corresponding to hard copy records in accordance with the existing QAPP.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Developing Open-Source Tools for Increasing the Efficiency of Synthetic Aviation Turbine Fuel Certification Process

FuelLib is an open-source Python-based fuel library, developed by NREL, that leverages the group contribution method (GCM) of [1] to systematically estimate the thermodynamic and transport properties of hydrocarbon fuels. FuelLib predicts these properties based on the molecular structure of individual compounds or compound families, using weight percentages of a fuel's composition, typically measured using techniques such as gas chromatography (GC). FuelLib enables property estimation over a wide range of temperatures and pressures of multi-component fuels in the absence of detailed molecular composition data, making it particularly valuable for complex fuel mixtures where detailed experimental characterization of fuel composition is unavailable. These capabilities contribute directly to synthetic aviation turbine fuels (SATF) development, supporting the short-term American Society for Testing and Materials (ASTM) qualification of drop-in fuels while potentially expanding ASTM boundaries to certify a broader range of fuels.

33 ADVANCED PROPULSION SYSTEMS↗

Laser Confocal Microscopy Uncertainty Quantification Study

At Los Alamos National Laboratory (LANL), the Storage Safety and Engineering (SSE) team completes annual surveillance on a subset of in-use interim nuclear material storage containers in fulfilment of requirements outlined in DOE Manual M 441.1-1. The containers are selected through several methods, such as subject matter expert judgement, random selection, and trending items. Following these selections, the SSE team has the capacity to complete surveillance on 15-20 containers each fiscal year, composed of a combination of SAVY-4000 and Hagan storage containers. Through previous work, the stainless-steel components of the containers have been identified as life limiting components, with an emphasis on the thin-walled bodies. The team is focused on understanding the extent of general and pitting corrosion, due to observations of extensive corrosion from stored contents and bag-out-bag degradation. Quantifying corrosion effects on the thin-walled stainless steel container bodies, and understanding potential impacts to the respective design release rates and design qualification release rates is paramount to the team. To date, destructive examination (DE) has proven to be the most insightful method for developing an understanding on the extent of corrosion on used containers. To standardize this process, the SSE team developed a destructive examination guide for analyzing stainless steel components of the containers. Corroded containers of interest are identified during surveillance activities and set aside for sectioning and characterization. Following sectioning, a major step in the DE workflow is the utilization of laser confocal microscopy for scanning corroded samples of interest and extracting data on pits, such as count, depth, and equivalent diameter. Adhering to the techniques outlined in the DE guide, analysis has been completed on two Hagans and one SAVY-4000 container, with the maximum pit depth recorded as 139.1 ± 22.82 μm on a 17.5 year old Hagan. The findings from the completed destructive examinations will be utilized to support lifetime extension efforts of the SAVY-4000 as the team can better estimate corrosion rates and effects over time based on stored contents and age. Due to the implications of observing extreme pit depths that approach the nominal container body thickness of .0299 inches (0.759 mm) or minimum container thickness of 0.236” (0.6 mm), high confidence in the LCM measurements is desired. Through testing outlined in, it was concluded that the total error ascribed to the 20x objective when conducting large image mapping on the Keyence VK-X3050 laser confocal microscope (LCM) relative to a 50x objective (reference) is 16.4% (± 8.73%). For shallow features on the order of pristine SAVY surface defects (i.e. 5 μm), this uncertainty is appropriate. However, this conservative estimate of total error poses a fundamental concern for pit depths that approach the thickness of the measured samples. That is, with the measurement uncertainty currently employed on all measurements, the LCM would be unable to resolve if a pit with a depth of 515 μm is through wall. Standard step height samples were procured and used in the present study to assess the resolution and repeatability of height measurements. Understanding the resolution and repeatability of height measurements was the first focus of the team as it relates directly to pit depth, which is of primary concern. Calibration gratings were procured to evaluate the resolution and repeatability of measurements in the X and Y axes of the LCM stage. The results of the depth uncertainty study were conducted first and presented in the subsequent sections. The planar uncertainty study is appended to the depth study with conclusions from both summarized at the end of the report.

36 MATERIALS SCIENCE↗

Advancing Industry 4.0: Multimodal Sensor Fusion for AI-Based Fault Detection in 3D Printing

Additive manufacturing, particularly fused deposition modeling, is transforming modern production by enabling rapid prototyping and complex part fabrication. However, its layer-by-layer process remains vulnerable to faults such as nozzle clogging, filament runout, and layer misalignment, which compromise print quality and reliability. Traditional inspection methods are costly, time-intensive, and often limited to post-process analysis, making them unsuitable for real-time intervention. In this current study, the authors developed a novel, low-cost, and portable faultdetection system that leverages multimodal sensor fusion and artificial intelligence for real-time monitoring in FDM-based 3D printing. The system integrates acoustic, vibration, and thermal sensing into a non-intrusive architecture, capturing complementary data streams that reflect both mechanical and process-related anomalies. Acoustic and thermal sensors operate in a fully contactless manner, while the vibration sensor requires minimal attachment such that it will not interfere with printer hardware, thereby preserving portability and ease of deployment. The multimodal signals are processed into spectrograms and time-frequency features, which are classified using convolutional neural networks for intelligent fault detection. The proposed system advances Industry 4.0 objectives by offering an affordable, scalable, and practical monitoring solution that improves faultdetection accuracy, reduces waste, and supports sustainable, adaptive manufacturing.

42 ENGINEERING↗

X-BASE: the first terrestrial carbon and water flux products from an extended data-driven scaling framework, FLUXCOM-X

Mapping in situ eddy covariance measurements of terrestrial land–atmosphere fluxes to the globe is a key method for diagnosing the Earth system from a data-driven perspective. We describe the first global products (called X-BASE) from a newly implemented upscaling framework, FLUXCOM-X, representing an advancement from the previous generation of FLUXCOM products in terms of flexibility and technical capabilities. The X-BASE products are comprised of estimates of CO 2 net ecosystem exchange (NEE), gross primary productivity (GPP), evapotranspiration (ET), and for the first time a novel, fully data-driven global transpiration product (ETT), at high spatial (0.05°) and temporal (hourly) resolution. X-BASE estimates the global NEE at −5.75 ± 0.33 Pg C yr −1 for the period 2001–2020, showing a much higher consistency with independent atmospheric carbon cycle constraints compared to the previous versions of FLUXCOM. The improvement of global NEE was likely only possible thanks to the international effort to increase the precision and consistency of eddy covariance collection and processing pipelines, as well as to the extension of the measurements to more site years resulting in a wider coverage of bioclimatic conditions. However, X-BASE global net ecosystem exchange shows a very low interannual variability, which is common to state-of-the-art data-driven flux products and remains a scientific challenge. With 125 ± 2.1 Pg C yr −1 for the same period, X-BASE GPP is slightly higher than previous FLUXCOM estimates, mostly in temperate and boreal areas. X-BASE evapotranspiration amounts to 74.7×10 3 ± 0.9×10 3 km 3 globally for the years 2001–2020 but exceeds precipitation in many dry areas, likely indicating overestimation in these regions. On average 57 % of evapotranspiration is estimated to be transpiration, in good agreement with isotope-based approaches, but higher than estimates from many land surface models. Despite considerable improvements to the previous upscaling products, many further opportunities for development exist. Pathways of exploration include methodological choices in the selection and processing of eddy covariance and satellite observations, their ingestion into the framework, and the configuration of machine learning methods. For this, the new FLUXCOM-X framework was specifically designed to have the necessary flexibility to experiment, diagnose, and converge to more accurate global flux estimates.

Nelson, Jacob A.↗

Hybrid Data‐Driven Discovery of High‐Performance Silver Selenide‐Based Thermoelectric Composites

Optimizing material compositions often enhances thermoelectric performances. However, the large selection of possible base elements and dopants results in a vast composition design space that is too large to systematically search using solely domain knowledge. To address this challenge, a hybrid data-driven strategy that integrates Bayesian optimization (BO) and Gaussian process regression (GPR) is proposed to optimize the composition of five elements (Ag, Se, S, Cu, and Te) in AgSe-based thermoelectric materials. Data is collected from the literature to provide prior knowledge for the initial GPR model, which is updated by actively collected experimental data during the iteration between BO and experiments. Within seven iterations, the optimized AgSe-based materials prepared using a simple high-throughput ink mixing and blade coating method deliver a high power factor of 2100 µW m −1 K −2 , which is a 75% improvement from the baseline composite (nominal composition of Ag 2 Se 1 ). In conclusion, the success of this study provides opportunities to generalize the demonstrated active machine learning technique to accelerate the development and optimization of a wide range of material systems with reduced experimental trials.

36 MATERIALS SCIENCE↗

A method for estimating light quenching in inorganic scintillator detectors for radioactive ion beam experiments

In recent experiments, inorganic scintillators have been used to study the decays of exotic nuclei, providing an alternative to silicon detectors and enabling measurements that were previously impossible. However, proper use of these materials requires us to understand and quantify the scintillation process, specifically in response to very heavy nuclei. Here, in this work, we show a simplified method based on the models of Birks (1951) and Meyer and Murray (1962) to parametrize the light output of inorganic scintillators in response to beams of energetic heavy ions over a broad range of energies. We test the accuracy of our parametrization approach by calculating light output and quenching factors for various ions and comparing them with experimental data from Lutetium Yttrium Orthosilicate (LYSO:Ce), a common inorganic scintillator. The Meyer–Murray model suggests that, for sufficiently heavy ions at high energies, the majority of the light output is associated with the creation of delta electrons, which are induced by the passage of the beam through the material. These delta electrons dramatically impact the response of detection systems when subject to ions with velocities typical of beams in modern fragmentation facilities. To illustrate this, we also present a qualitative estimate of the effects of delta rays on overall light output using the Birks–Meyer–Murray parametrization. The approach presented herein will serve as a basic framework for further, more rigorous studies of scintillator response to heavy ions. This work is a crucial first step in planning future experiments where energetic exotic nuclei are interacting with scintillator detectors.

Heavy ion↗

Novel Hot Gas Components for Gas Turbine Engines Enabled by Materials and Additive Manufacturing Process Development

Additive Manufacturing (AM), also known as 3D printing, has emerged as a manufacturing method that enables new design freedom for gas turbine engine manufacturers. However, the material selection for AM processable high-temperature super alloys is currently limited. Additionally, the heat transfer performance of AM enabled micro-cooling architectures is not yet well understood. Accordingly, in support of advanced manufacturing and engine performance development, Oak Ridge National Laboratory (ORNL)and Solar Turbines (Solar) conducted a multidisciplinary project to generate both AM super alloy material properties data and micro-channel performance data for two AM super alloys. The data supported the design and analysis of an internally cooled turbine hot section AM tip shoe component. This data was used to analytically predict the reduction in operating temperature of a gas turbine tip shoe. The work concluded that the cooling flow required to cool the tip shoe can be tuned to suit the efficiency improvements desired in an industrial gas turbine.

36 MATERIALS SCIENCE↗

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]↗

Deep learning-driven super-resolution in Raman hyperspectral imaging: Efficient high-resolution reconstruction from low-resolution data

Deep learning (DL) has become an indispensable tool in hyperspectral data analysis, automatically extracting valuable features from complex, high-dimensional datasets. Super-resolution reconstruction, an essential aspect of hyperspectral data, involves enhancing spatial resolution, particularly relevant to low-resolution hyperspectral data. Yet, the pursuit of super-resolution in hyperspectral analysis is fraught with challenges, including acquiring ground truth high-resolution data for training, generalization, and scalability. The pressing issue of extended spectral acquisition times, notably for high-resolution scans, is a significant roadblock in hyperspectral imaging. Super-resolution methods offer a promising solution by providing higher spatial resolution data to expedite data collection and yield more efficient outcomes. This paper delves into a practical application of these concepts using Raman imaging, where spectral acquisition times can be prohibitively long. In this context, DL-based super-resolution models demonstrate their efficacy by predicting and reconstructing high-resolution Raman data from low-resolution input, eliminating the need for resource-intensive high-resolution scans. While previous work often relied on substantial high-resolution datasets, this study showcases the ability to achieve similar outcomes even with limited data, presenting a more practical and cost-effective approach. In conclusion, the results offer a glimpse into the transformative potential of this technology to streamline hyperspectral imaging applications by saving valuable time and resources through the successful generation of high-resolution data from low-resolution inputs.

42 ENGINEERING↗

Decoding diffraction and spectroscopy data with machine learning: A tutorial

This Tutorial provides a step-by-step guide on how to apply supervised machine-learning techniques to analyze diffraction and spectroscopy data. This Tutorial details four models—a reconstruction-focused model, a regression-focused model, a hybrid reconstruction/regression model, and a multimodal model—that use x-ray diffraction profiles and vibrational density of states spectra to predict various microstructural descriptors. In this Tutorial, we cover data pre-processing steps, constructions of the models via dimensionality reduction and regression, training, and analysis of these models. Comparisons of the model’s performance are provided, highlighting the strength and weakness of the various approaches utilized.

36 MATERIALS SCIENCE↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗

An AI-accelerated pathway for reproducible and stable halide perovskites

Halide perovskites (HPs) have remarkable optoelectronic properties, and in the last decade their photovoltaic power conversion efficiency and light-emitting diode efficiency have skyrocketed. Despite the surge in research on these burgeoning materials, two key challenges in the field remain: material irreproducibility and instability. Their behavior is especially dynamic in response to environmental stressors, due to complex interactions with the perovskite crystal lattice. Here, in this review, we survey the latest achievements in HP materials research accomplished with the assistance of artificial intelligence (AI), through the implementation of automated experimentation and machine learning (ML) data analysis. Automated synthesis and characterization tackle problems with material irreproducibility by systematically controlling parameters with very high precision, creating massive datasets, and allowing methodical comparisons from which unbiased conclusions can be drawn. AI can reveal otherwise unnoticed trends, inform future experiments with the highest potential information gain, and forecast future performance. The review concludes with a forward viewpoint of how human-assisted closed-loop laboratories and shared databases allow halide perovskite materials’ processing, properties, and performance to be potentially optimized with AI, accelerating the development of highly reproducible and stable optoelectronic devices.

Hering, Abigail R. [Univ. of California, Davis, CA↗