Search NASA⌕ Search

SEARCH · Search NASA

Results for “Feature selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Characterization of an unusual SARS-CoV-2 main protease natural variant exhibiting resistance to nirmatrelvir and ensitrelvir

We investigate the effects of two naturally selected substitution and deletion (Δ) mutations, constituting part of the substrate binding subsites S2 and S4, on the structure, function, and inhibition of SARS CoV-2 main protease. Comparable to wild-type, MPro D48Y/ΔP168 undergoes N-terminal autoprocessing essential for stable dimer formation and mature-like catalytic activity. The structures are similar, but for an open active site conformation in MPro D48Y/ΔP168 and increased dynamics of the S2 helix, S5 loop, and the helical domain. Some dimer interface contacts exhibit shorter H bond distances corroborating the ~40-fold enhanced dimerization of the mutant although its thermal sensitivity to unfolding is 8 °C lower, relative to wild-type. ITC reveals a 3- and 5-fold decrease in binding affinity for nirmatrelvir and ensitrelvir, respectively, and similar GC373 affinity, to MPro D48Y/ΔP168 relative to wild-type. Structural differences in four inhibitor complexes of MPro D48Y/ΔP168 compared to wild-type are described. Consistent with enhanced dynamics, the S2 helix and S5 loop adopting a more open conformation appears to be a unique feature of MPro D48Y/ΔP168 both in the inhibitor-free and bound states. Our results suggest that mutational effects are compensated by changes in the conformational dynamics and thereby modulate N-terminal autoprocessing, K dimer , catalytic efficiency, and inhibitor binding.

60 APPLIED LIFE SCIENCES↗

Symmetry breaking of fluorophore binding to a G-quadruplex generates an RNA aptamer with picomolar K D

Fluorogenic RNA aptamer tags with high affinity enable RNA purification and imaging. The G-quadruplex (G4) based Mango (M) series of aptamers were selected to bind a thiazole orange based (TO1-Biotin) ligand. Using a chemical biology and reselection approach, we have produced a MII.2 aptamer–ligand complex with a remarkable set of properties: Its unprecedented K D of 45 pM, formaldehyde resistance (8% v/v), temperature stability and ligand photo-recycling properties are all unusual to find simultaneously within a small RNA tag. Crystal structures demonstrate how MII.2, which differs from MII by a single A23U mutation, and modification of the TO1-Biotin ligand to TO1-6A-Biotin achieves these results. MII binds TO1-Biotin heterogeneously via a G4 surface that is surrounded by a stadium of five adenosines. Breaking this pseudo-rotational symmetry results in a highly cooperative and homogeneous ligand binding pocket: A22 of the G4 stadium stacks on the G4 binding surface while the TO1-6A-Biotin ligand completely fills the remaining three quadrants of the G4 ligand binding face. Similar optimization attempts with MIII.1, which already binds TO1-Biotin in a homogeneous manner, did not produce such marked improvements. We use the novel features of the MII.2 complex to demonstrate a powerful optically-based RNA purification system.

59 BASIC BIOLOGICAL SCIENCES↗

An In Situ Feed Monitoring System for Molten Salt Reactors with Fast Neutron Energy Spectrum Molten Salt Reactor Applications

Molten salt reactors (MSRs) are one of the six promising advanced reactor technologies selected for further research and development by the Generation IV International Forum. More than twenty MSR designs are actively being developed around the world. Several of these designs are liquid-fueled and intended for operation within the fast neutron energy spectrum.1 National regulations will require liquid-fueled MSRs to control and account for nuclear material within licensed facilities. Additionally, states with comprehensive safeguards agreements with the International Atomic Energy Agency (IAEA) are obligated to declare nuclear material quantities within facilities. In return, the IAEA Department of Safeguards independently verifies these quantities and provides assurance that the nuclear material and facility are being used only for peaceful purposes. One key distinction of liquid-fueled MSRs compared with other types of reactors is that in portions of the facility, the nuclear material is in bulk form rather than discrete items. Traditional nuclear material accounting techniques such as physical item counting and verification of serial numbers on fresh fuel assemblies do not translate directly to all process streams within liquid-fueled reactors. Liquid-fueled MSRs are typically designed with low excess reactivity. This feature provides safety benefits but also means that most MSRs require the addition of makeup fuel salt while a reactor is operational. The nuclear material in the initial fuel salt and in any makeup fuel salt must be quantified. Additionally, distinct nuclear material diversion and reactor misuse scenarios form the basis of the detection methods and monitoring systems developed for liquid-fueled MSRs. For example, the IAEA provides assurance that fuel salt containing nuclear material is not being diverted from the system, that the feed salt matches the reported actinide concentrations and uranium enrichment, and that no additional fertile material is being introduced into the system. Measurement systems currently used for nuclear material control and accounting are not directly applicable to achieving MSR safeguards goals. This paper concerns a system being designed to account for the nuclear material added to liquid-fueled MSRs and monitor for diversion and misuse scenarios related to MSR feed systems.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Dynamic Features of Cu-Ceria Interface under CO 2 Hydrogenation to Methanol

It is generally accepted that metal–support interaction is very important for the hydrogenation of CO 2 to methanol, but little has been revealed about the feature of interfacial active sites under real reaction conditions since there are only limited techniques that can be applied under high-pressure conditions. Here, in this work, by combining multiple in situ and operando techniques on a model Cu/ceria catalyst, we have tracked Cu and ceria sites for methanol formation. Under the reaction condition, it is found that upon reaching the reaction temperature, oxidized Cu species in the as-synthesized catalyst immediately change into metallic Cu species. Following this, it is the gradual formation of methanol, the changing rate of which coincides with the formation of a unique Ce 3+ species. The combined experimental results and density functional theory (DFT) calculations have determined that the formed Ce 3+ sites driven by the reaction conditions are bound to hydrides, adsorbed carbonate species, and interfacial active Cu sites. The Cu-ceria interaction in this complex moiety is weak and can be easily disturbed with reaction environment variations, leading to dynamic changes at the interface upon the hydrogenation of active carbonate intermediates, which are precursors for the formation of methanol. The formation of this unique Cu–Ce 3+ interface and its dynamicity lead to an increase of methanol selectivity from less than 20% to 60%. These results suggest that reactant-derived species (H – and carbonate in this work) can be essential components of the active center with the functions of manipulating the metal−oxide interaction and directing reaction pathways.

CO2 hydrogenation to methanol↗

Asymmetric jet shapes with two-dimensional jet tomography

Two-dimensional (2D) jet tomography is a promising tool to study jet medium modification in high-energy heavy-ion collisions. It combines gradient (transverse) and longitudinal jet tomography for selection of events with localized initial jet production positions. It exploits the transverse asymmetry and energy loss that depend, respectively, on the transverse gradient and jet path length inside the quark-gluon plasma (QGP). In this study, we employ 2D jet tomography to study medium modification of the jet shape of γ -triggered jets within the linear Boltzmann transport (LBT) model for jet propagation in heavy-ion collisions. Our results show that jets with small transverse asymmetry ( A N n ⃗ ) or small γ -jet asymmetry ( x J γ = p T jet / p T γ ) exhibit a broader jet shape than those with larger A N n ⃗ or x J γ , since the former are produced at the center and go through longer path lengths while the later are off center and close to the surface of the QGP fireball. In events with finite values of A N n ⃗ , jet shapes are asymmetric with respect to the event plane. Hard partons at the core of the jet are deflected away from the denser region while soft partons from the medium response at large angles flow toward the denser part of QGP. Future experimental measurements of these asymmetric features of the jet shape can be used to study the transport properties of jets and medium responses. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC↗

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗

Molecular Mass Growth Processes to Polycyclic Aromatic Hydrocarbons through Radical–Radical Reactions Exploiting Photoionization Reflectron Time-of-Flight Mass Spectrometry

Polycyclic aromatic hydrocarbons (PAHs) represent critical building blocks in molecular mass growth processes to carbonaceous nanoparticles, referred to as interstellar and circumstellar grains along with soot particles in astrophysical environments and combustion systems, respectively. Recent advancements on elucidating elementary steps to PAHs have utilized reactions of aromatic radicals, resonantly stabilized free radicals, and aliphatic radicals with closed shell hydrocarbons. However, the role of radical–radical reactions (RRRs) leading to PAHs has remained largely unexplored on the molecular level due to preceding experimental challenges in producing sufficiently high number densities of radical reactants for isomer-selective detection of products from bimolecular and termolecular reactions. This Account offers the latest developments in our knowledge on the mechanisms and pathways to PAHs via RRRs probed in a chemical microreactor at temperatures as high as 1600 K. Product preservation in a molecular beam coupled with synchrotron vacuum ultraviolet photoionization reflectron time-of-flight mass spectrometry and photoelectron photoion coincidence spectroscopy enabled isomer-selective detection of PAHs of up to three rings by their photoionization efficiency curves, which were fit with a linear combination of reference curves for identification. Experiments were combined with computational fluid dynamics modeling of the physicochemical processes in the microreactor, as well as high-level electronic structure calculations to reveal the reaction pathways of each system. Six distinct reaction mechanisms were discovered in this work: propargyl addition─benzannulation (PABA), methyl addition─ring expansion (MARE), cyclopentadienyl addition─naphthylization (CPAN), fulvenallenyl addition─cyclization─aromatization (FACA), benzyl addition─aromatization (BAA), and phenyl addition─pentacyclization (PAP). By systematically varying the number of carbon atoms in the radical reactants, molecular mass growth processes involving reactions between radicals with odd numbers of carbon atoms access aromatics carrying one, two, or three six-membered rings, whereas reactions between even- and odd-carbon-numbered radicals produce aromatics combining five- and six-membered rings. Our investigations reveal unconventional cycloadditions on excited state triplet surfaces, additions of radicals to low spin density carbon-centered radicals, spiroaromatic and fulvene-type intermediates, and highly strained bicyclic reaction intermediates, challenging current perceptions of PAH molecular mass growth processes. All of the listed mechanisms, except for FACA, feature endoergic reactions or barriers which lie above the separated reactants and therefore might be central to circumstellar environments of carbon-rich stars and planetary nebulae as their descendants, but they play no role in the gas phase of cold molecular clouds where temperatures as low as 10 K dominate. Altogether, this work provides detailed reaction mechanisms of PAH growth processes, advancing our knowledge of the chemistry of carbonaceous matter in the universe.

Addition reactions↗

Suppressed paramagnetism in amorphous Ta 2 O 5 − x oxides and its link to superconducting-qubit performance

Amorphous-oxide layers in thin-film capacitors are linked to reduced transmon-qubit T 1 coherence times. Ta -based capacitors outperform Nb -based ones, suggesting that amorphous Ta 2 O 5 − x is less lossy than Nb 2 O 5 − x . We investigate the microscopic features of these amorphous oxides using ab initio molecular dynamics and density functional theory, revealing the origins of the superior performance of Ta 2 O 5 − x . We establish that oxygen deficiency is less likely to occur in amorphous Ta 2 O 5 − x than in Nb 2 O 5 − x for 0 ≤ x ≤ 0.25 and that for a given oxygen deficiency x , metal Ta — Ta bond formation is enhanced. Such bonds, which are accommodated by structural flaws in the amorphous network, capture electrons better than in amorphous Nb 2 O 5 − x . These thermochemical differences quench or highly suppress magnetic moments in amorphous Ta 2 O 5 − x and eliminate a potential source of quasiparticles and magnetic flux noise. We also show that hyperfine couplings between Nb nuclei and local magnetic moments in Nb 2 O 5 − x can form “two-level systems” (TLSs) or “two-level fluctuators” with energy splittings of 100–1000 MHz or higher. This reveals a TLS mechanism in amorphous Nb 2 O 5 − x oxide layers that is likely inactive in Ta 2 O 5 − x . Our work provides a fundamental understanding of the materials chemistry and limitations imposed by native oxides of superconducting qubits that can be used to guide materials selection and processing.

Pritchard, P. Graham [Northwestern U.] (ORCID:0000↗

Tip-based proximity ferroelectric switching and piezoelectric response in wurtzite multilayers

Proximity ferroelectricity is a paradigm for inducing ferroelectricity when a nonferroelectric polar material (such as Al⁢ N), which is unswitchable with an external field below the dielectric breakdown field, becomes a practically switchable ferroelectric in direct contact with a thin switchable ferroelectric layer (such as Al 1−𝑥 ⁢Sc 𝑥 ⁢N). Here, we develop a Landau-Ginzburg-Devonshire approach to study the proximity effect of local piezoelectric response and polarization reversal in wurtzite ferroelectric multilayers under a sharp electrically biased tip. Using finite-element modeling, we analyze the probe-induced nucleation of nanodomains, the features of local polarization hysteresis loops and coercive fields in the Al 1−𝑥 ⁢Sc 𝑥 ⁢N/Al⁢ N bilayers and three-layers. Similar to the wurtzite multilayers sandwiched between two parallel electrodes, the regimes of “proximity switching” (when all layers collectively switch) and the regime of “proximity suppression” (when they collectively do not switch) are the only two possible regimes in the probe-electrode geometry. However, the parameters and asymmetry of the local piezoresponse and polarization hysteresis loops depend significantly on the sequence of the layers with respect to the probe. The physical mechanism of proximity ferroelectricity in the local probe geometry is a depolarizing electric field determined by the polarization of the layers and their relative thickness. The field, whose direction is opposite to the polarization vector in the layer(s) with the larger spontaneous polarization (such as Al⁢ N), renormalizes the double-well ferroelectric potential to lower the steepness of the switching barrier in the “otherwise unswitchable” polar layers. Tip-based control of domains in otherwise nonferroelectric layers using proximity ferroelectricity can provide nanoscale control of domain reversal in memory, actuation, sensing, and optical applications. The ability of the tip-induced proximity switching to differentially switch multilayers, based on the order of the layers, provides a powerful tool for selective domain engineering.

36 MATERIALS SCIENCE↗

Multimodal Nanoscale Mapping of Local Structure and CO 2 Adsorption in Metal–Organic Frameworks

Diamine functionalization of the metal−organic framework Mg 2 (dobpdc) (dobpdc 4− = 4,4′-dioxidobiphenyl-3,3′-dicarboxylate) significantly enhances its selectivity for CO 2 capture from flue gases and air. The structure and CO 2 capacity of such materials are typically assessed using bulk techniques that rely on averaging signal over large ensembles of unit cells, obscuring local heterogeneities, such as variations in CO 2 occupancy across individual nanocrystals. To resolve this limitation, we demonstrate a multimodal, nanoscale characterization of Mg 2 (dobpdc) appended with 1,3-diaminopropane. By employing recently developed characterization techniques at progressively smaller length scales, we uncover insights from correspondingly smaller populations of unit cells. First, we use parallel-beam 3D electron diffraction (3D ED) to identify a prominent expansion in lattice parameters upon desorption of CO 2 , as observed at the level of single nanocrystals. Second, we use convergent-probe 4D scanning transmission electron microscopy (4D-STEM) to quantify associated differences in lattice strain as a function of gas loading and diamine appending. These measurements sample small subvolumes within individual nanocrystals. Finally, we apply infrared scattering scanning near-field optical microscopy (IR s- SNOM) to confirm variable CO 2 chemisorption across adsorption sites at the surface of single nanocrystals. This multimodal, multiscale approach allows us to map heterogeneity within individual nanocrystals. Collectively, these findings emphasize the importance of local, nanoscale characterization of metal−organic frameworks in revealing previously unresolvable features that impact their performance.

Karstens, Sarah L. [University of California, Berk↗

Multi-plane moment-of-fluid interface reconstruction in 3D

Moment-of-fluid (MOF) methods for interface reconstruction approximate the region occupied by material in each mesh element only through reference to its geometric moments. Here, we present a 3D MOF method that represents the material (POM) in each cell as the convex intersection of the cell and multiple half-spaces, each selected to minimize the least-squares error between computed moments of the approximated material and provided reference moments. This optimization problem is highly non-linear and non-convex, making the numerical result very sensitive to the initial guess. To create an effective initial guess in each cell, we construct an ellipsoid from 0th–2nd order reference moments such that its shape corresponds with that of the POM. Within this ellipsoid we inscribe a polyhedron, and initialize the minimization problem with the half-spaces defined by each of its faces. The inscribed polyhedron has minimally 4 faces, and using up to 3rd order moments permits optimization over up to 20 unknown values. We therefore define MOF methods that utilize 4, 5, or 6 half-spaces, correspondingly initialized with the faces of a single inscribed tetrahedron, triangular prism, or hexahedron. Stability of the non-linear optimization is further improved with a prepossessing step that normalizes the reference moments according to the axes of the reference ellipsoid. Using this approach, the non-linear least-squares solver reliably converges to a near-global minimum from a single initial guess. We demonstrate accuracy and robustness using single-cell and multi-cell examples over a wide spectrum of geometry. In particular, we demonstrate our ability to exactly reproduce several important and complex features defined by up to four half-spaces, such as corners, filaments, filament tips, and embedded material in the cell.

3D interface reconstruction↗

Statistical evaluation of microscale stress conditions leading to void nucleation in the weak shock regime

Here, we investigate the heterogeneity of the stress state driven by anisotropic deformation response at the single crystal level through five statistical volume element (SVE) calculations of polycrystalline BCC tantalum. This work focuses on grain boundaries as a prominent material defect type prone to void nucleation based upon experimental observations of predominantly intergranular void nucleation in this material. The SVEs are constructed to be statistically representative of larger volumes of material and are meshed such that mean and standard deviation of grain size and orientation information is reconstructed. The computational meshes feature hexahedral (brick) elements and smooth conformal grain boundaries where significant stress concentration is known to occur, a tail effect of interest in the extreme events process of dynamic ductile damage. An existing micromechanical crystallographic plasticity model shown to capture the single crystal behavior of BCC tantalum well is used to perform the polycrystal calculations. The model includes representation of the non-Schmid effect of non-planar screw dislocation kinetics in tantalum. A three-dimensional stress state time profile predicted by damage modeling of a flyer plate impact experiment is applied as boundary conditions to each SVE. Resulting grain boundary stress state statistics are strongly non-Gaussian. Significant structural evolution is observed within the compressive hold before unloading into tension in the stress profile. Strong angular dependence of grain boundary traction magnitude with shock direction is observed. Non-Schmid effects continue to suggest their influence on propensity of microstructural defect types to nucleate voids. A general void nucleation criterion is proposed using probability theory. The general framework is specified to polycrystalline BCC tantalum in the weak shock regime to include the SVE calculations and literature molecular dynamics calculations of grain boundary void nucleation strength. Probability density functions (PDFs) are used to describe the interaction between the local stress state heterogeneity and the distributed grain boundary void nucleation strength state. A causation entropy maximization procedure removes the requirement for ad hoc selection of a PDF functional form and provides a rigorous procedure for data-based PDF determination. The resulting physically informed PDF describes the spatial appearance frequency of nucleated voids as a function of applied macroscale pressure. Lower length scale physics are thus packaged in a precise and computationally efficient way to provide computational plasticity insight to macroscale dynamic ductile damage models.

36 MATERIALS SCIENCE↗

Geology of the One Earth Energy Site

The One Earth Energy site is one of two sites in the Illinois Storage Corridor (ISC) project. The objectives of the ISC project is to accelerate commercial deployment of carbon capture utilization and storage at two individual sites and receive approvals for Underground Injection Control (UIC) Class VI permits for construction at each site. At the One Earth Energy site, an extensive data collection program was undertaken, which included the drilling of a test well (One Earth Energy #1 [OEE #1]), four 2D seismic lines, and a small 3D seismic survey. The OEE #1 well was drilled in 2022 and acquired extensive core, log, and testing data to characterize the subsurface geology of the site. Coring was focused on the storage interval, the Mt. Simon Sandstone, and the confining interval, the Eau Claire Formation. The core and log data were used to evaluate the sedimentology and sequence stratigraphy, as well as to develop the conceptual geologic model. This report includes the geological summaries of the Mt. Simon Sandstone and the Eau Claire Formation. The extensive analysis of the log data is included in the petrophysical section, showing ranges of porosity, estimated pore size, and the mineral content of selected zones in the well. The separate petrographic technical report entitled “Petrographic and Advanced Geologic Characterization Report on One Earth Energy #1 (API# 1211325373)”, report number DOE-UIUC-0031892-04, details thin section point-counting analysis that includes mineralogical and pore space analysis, including grain size analysis, annotated thin section photomicrographs, scanning electron microscopy (SEM) with energy dispersive X-ray spectroscopy (EDS), and statistics of grain size analysis on Mt. Simon thin sections from OEE #1. The final OEE #1 well data to be included in this geology report is the routine core analysis of both whole core plugs and rotary sidewall core plugs. In addition to the OEE #1 well, four 2D seismic lines and a small 3D survey were acquired as part of the overall subsurface geological characterization. This geology report references the seismic interpretation report, entitled “One Earth Energy Site Seismic Interpretation Task 5.0”, report number DOE-UIUC-0031892-07. This report details the stratigraphic and structural interpretation of the 2D and 3D seismic data acquired at the One Earth Energy site. The 2D seismic data was acquired in 2019 and 2021, and the 3D survey was acquired in 2022. The objectives of the seismic programs were to contribute to the subsurface characterization of the Mt. Simon-Eau Claire Storage Complex by evaluating the continuity of potential storage reservoirs and containment intervals across the project area, and to determine if any geologic features are present that would increase containment risk to the proposed carbon storage project.

09 BIOMASS FUELS↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

Exceptional Electrical Detection of Trace NO 2 via Mixed Metal MOF-on-MOF Film-Based Sensors

The tunability of metal–organic frameworks (MOFs) makes them exceptional materials for the development of highly selective, low-power sensors for toxic gas detection. Herein, we demonstrate enhanced detection of NO 2 gas by a MOF-based electrical impedance sensor made using a unique mixed metal MOF-on-MOF synthesis. For this work, a combined experimental and computational study was performed using the exemplar Ni x Mg 1–x -MOF-74 to understand the fundamental structure–property relationships behind metal mixing and MOF film synthesis methods on sensor performance. Density functional theory results indicated that the presence of Ni in Mg-MOF-74 increased framework stability and increased the electron density of states at lower energies near the HOMO, as well as enhanced the NO 2 –Mg adsorption interaction. Impedance data of the Ni x Mg 1–x -MOF-74 films with larger Ni contents showed greater impedance change after exposure to 1 ppm of NO 2 gas. Furthermore, when synthesized through either a drop-cast or direct solvothermal film growth approach, the monometallic Ni-based sensors had the best performance. However, the mixed metal Ni x Mg 1–x -MOF-74 sensors synthesized through a MOF-on-MOF approach resulted in the highest impedance change, outperforming all monometallic Ni-based sensors. In particular, the mixed metal Ni-on-Mg-MOF-74 film was the best-performing sensor with an impedance change of 309 upon trace NO 2 exposure. Change in impedance response after NO 2 exposure was improved by 52% compared to the best monometallic Ni-on-Ni-MOF-74 sensor. Structural analysis of the Ni-on-Mg film showed that the first Mg-MOF-74 layer acts as a structural template controlling the structural features of the final film after metal exchange with Ni. This led to improved film quality, evidenced by the greater crystallinity and larger MOF grain sizes, and resulted in enhanced sensor performance which was not achievable through other metal mixing methods. Altogether, this study identifies structure–property relationships and synthetic templating methods that inform MOF-based sensor design, allowing for improved detection of toxic compounds.

36 MATERIALS SCIENCE↗

Idaho National Laboratory Integrated Multisite SSHAC Level 3: Probabilistic Volcanic Hazards Assessment

The Idaho National Laboratory (INL) resides on the eastern Snake River Plain (ESRP), part of the Snake River Plain (SRP) with a complex origin and geologic history of volcanism. Much of the Quaternary (last 2.58 million years) volcanism within 400 km of INL is genetically associated with a major thermal anomaly referred to as the Yellowstone hotspot, which is currently located more than 180 km northeast of INL. Near INL, local volcanic sources include silicic domes near its southern border, and numerous dike-fed basaltic vents of the ESRP, some of which are near or within the INL boundaries. Quantitative probabilistic assessments of screened-in volcanic hazardous phenomena for nine different facility complexes at the INL are presented for a Senior Seismic Hazard Analysis Committee (SSHAC) Level 3 (SL3) study. The comprehensive, integrated multisite SSHAC study consists of a single regional Probabilistic Volcanic Hazards Assessment (PVHA) that pertains to all nine INL facility complexes, with site-specific information developed at each respective area of interest (AOI), referred to as the "INL facility AOI" (Figure ES-1). Hazard products are generated for specified INL facility AOIs for use by multiple stakeholders from different agencies in risk-informed decision-making regarding site selection, operations, and design of nuclear facilities at INL, consistent with U.S. Department of Energy (DOE) and U.S. Nuclear Regulatory Commission (NRC) regulatory guidance. The study also serves as the basis for future periodic safety assessments required for existing DOE facilities, such as 10-year evaluations of natural-phenomena hazards. Elements of the INL PVHA including initial characterization, screening, quantitative assessments of eruption and hazard potential, and consideration of facility needs are conducted using the three phases of the SSHAC process: evaluation, integration, and documentation. As per regulatory guidance, the SSHAC framework provided the necessary processes and procedures for the PVHA Technical Integration (TI) team, at three workshops, five formal working meetings and many TI team remote meetings, to conduct initial characterization, screen the volcanic hazards, create PVHA model inputs, exercise those models in the PVHA, consider facility-specific volcanic hazard needs, and perform final hazard calculations. The Participatory Peer Review Panel (PPRP) provided independent oversight and performed process and technical reviews of the PVHA throughout its duration. The study included an extensive New Data Collection and Analyses (NDCA) program developed by the PVHA TI team to reduce uncertainties in hazard-significant elements in the PVHA model. NDCA activities generated 20 reports providing important contributory datasets and results to the project database for characterizing the ESRP. For example, a report compiling the dimensions of ESRP shield volcanoes and lava fields was used to construct volcanic footprints (areas of impact), was compared with data from INL subsurface cores, and was used to validate the results of lava-flow inundation modeling on the contemporary terrain. Another example is the acquisition of aeromagnetic data over INL and its surrounding area, with maps of buried magmatic features (e.g., subsurface volcanoes and swarms of feeder dikes) that informed the PVHA conceptual model of volcanism. The SSHAC evaluation process was used to conduct all elements of the PVHA including initial characterization and screening. Existing data and NDCA activities provided the foundation to develop the tectonomagmatic conceptual model of volcanism for the region of geographical interest in the SRP and Yellowstone hotspot volcanic system, and for considering volcanoes in the western US. Considering Quaternary volcanic sources active during this period, the screening approach identified and evaluated magma compositions, types of eruptive and intrusive phenomena, types of hazardous phenomena, and proximity of sources to INL. The PVHA TI team evaluated 18 types of magmatic sources in terms of 20 potentially hazardous phenomena, resulting in 360 screening decisions. The screening process resulted in 140 screened-in hazardous volcanic phenomena for Quaternary volcanic sources 1) proximal to INL facility complexes in the ESRP (64 basaltic and 54 silicic), 2) regional sources associated with the Yellowstone caldera system and Blackfoot Reservoir volcanic field (7), and 3) more distal sources from thirteen Cascade volcanoes and two volcanoes at Long Valley caldera (CA).

58 GEOSCIENCES↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗