Search NASA⌕ Search

SEARCH · Search NASA

Results for “functional convergence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Modeling and Simulation of Fuel Dispersal During the Loss-of-Coolant Accident

This document is the compilation of the milestone portion to a larger end of project NEUP report. The executive summary of the modeling portion is provided below: In the event of cladding rupture during a postulated LOCA in a pressurized water reactor, fuel particles, along with fission gases, can be expelled into the reactor core from the fractured fuel rod, a phenomenon referred to as fuel dispersal. The initial stage of fuel dispersal is strongly influenced by the high-pressure ejection of fuel fragments, the size and geometry of the ruptured cladding, and the depressurization history of the fuel rod during the postulated LOCA transient. Depending on the location of the burst orifice relative to the quench front, the dispersal event represents an intricate three-phase flow and heat transfer phenomenon, where high-temperature fuel particles carried by the fission gases interact with the coolant within the narrow subchannels of the fuel assemblies, inducing localized phase change. Given the unique multiphysics nature of this phenomena, the current study develops a dedicated computational framework to predict the mass distribution and cooling of dispersing fuel particles, facilitating post-accident assessment and management of the fuel assemblies. Considering the scale of nuclear reactor applications, a continuum three-fluid model is proposed for simulating the transport of solids within the reactor core. With high-temperature fuel fragments within the liquid media, nucleation sites inducing phase changes are dispersed within the flow domain. Coupled with the fact that the transient dispersal event occurs on different time scales than other three-phase flow applications, this study derives a time-averaged three-fluid flow model without losing generality. The assumptions regarding the continuum treatment of the solid phase and the modeling of fuel dispersal behavior are incorporated to simplify the governing equations and derive applicable closure relations. The computational validation of the model was conducted using adiabatic experimental results obtained from ongoing research at Oregon State University, focusing on characterizing fuel dispersal behavior during simulated LOCA conditions. Settlement characteristics of the solids, quantified by the probability distribution of equivalent particles, closely matched the probability density functions reported in experimental studies. The transport of fuel particles within a scaled 5 × 5 lattice of a pressurized-water reactor rod bundle geometry was modeled through a two-fluid Eulerian framework. The required boundary conditions were evaluated from the fuel performance code BISON in a postulated large-break LOCA scenario. The modeling framework considered solid fuel particles as granular matter, interacting with the gaseous dry steam phase and fission gases through the governing interfacial momentum exchange between the participating fluids. The simulation results provided the volume fraction of the solids obtained at the bottom surface of the enclosing tank geometry. Postulated LOCA leading to fuel dispersal phenomena involves the strong coupling between fuel thermomechanics, cladding deformation, thermal-hydraulics, and fuel particle transport. Incorporation of such a strong coupling in numerical simulation is performed by coupling the multiphysics solvers. In the case of fuel dispersal, a strong coupled simulation can be performed by coupling the BISON code for fuel performance, the TRACE code for system-level thermal hydraulics, and fuel particle transport in Multiphysics Object-Oriented Simulation Environment (MOOSE). For such intricate infrastructure, the MOOSE Framework eases the data transfer between codes. The recent version of MOOSE has incorporated the Navier-Stokes module for the fluid flow. An exploratory exercise was done to gain familiarity with finite volume capabilities in the MOOSE framework to incorporate the Spalart-Allmaras (SA) turbulence model. New finite-volume and auxiliary kernels were introduced to assemble the SA transport equation, compute turbulent viscosity, and evaluate wall distance and diagnostic turbulence terms, fully integrated with existing Navier-Stokes modules. A turbulent lid-driven cavity at a Reynolds number of approximately 10,000 is used for verification. MOOSE shows the robust solver convergence and produces the turbulent features. But it underpredicts the velocity profile and turbulent quantities, emphasizing the need to develop improved SA near-wall treatments (e.g., low-Re corrections or wall functions) as a key direction for future work.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Convergent Concordant Mode Approach for Molecular Vibrations: CMA-2

The concordant mode approach (CMA) is a promising new scheme for dramatically increasing the system size and level of theory achievable in quantum chemical computations of molecular vibrational frequencies. Here, we achieve advances in the CMA hierarchy by computations targeting CCSD(T)/cc-pVTZ (coupled cluster singles and doubles with perturbative triples using a correlation-consistent polarized-valence triple-ζ basis set) benchmarks within the G2 molecular test set, executing a statistical analysis for 1501 frequencies from 111 compounds and then separately solving the refractory case of pyridine. First, MP2/cc-pVTZ (second-order Møller–Plesset perturbation theory with the same basis set) proves to be an excellent and preferred choice for generating the underlying (Level B) normal modes of the CMA scheme. Utilizing this Level B within the CMA-0A method reproduces the 1501 benchmark frequencies with a mean absolute error (MAE) of only 0.11 cm –1 and an attendant standard deviation of 0.49 cm –1 . Second, a convergent CMA-2 method is constituted that allows efficient computation of higher level (Level A) frequencies to any reasonable accuracy threshold by using only Hartree–Fock (HF) and MP2 or density functional theory (DFT) data to generate ξ parameters, which select the sparse off-diagonal force field elements for explicit evaluation at Level A. When Level B = MP2/cc-pVTZ, a cutoff of ξ = 0.02 provides an average maximum absolute error per molecule of only 0.17 cm –1 by incurring merely a 33% increase in average cost over CMA-0A. This CMA-2 method also eradicates the 4 problematic CMA-0A outliers of pyridine with even less effort (ξ = 0.04, 22% increase). Finally, the newly developed CMA procedures are shown to be highly successful when applied to 1-(1H-pyrrol-3-yl)ethanol, a new test molecule with diverse types of vibration.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrating Intermediate Traits in Phylogenetic Genotype-to-Phenotype Studies

A major goal of research in evolution and genetics is linking genotype to phenotype. This work could be direct, such as determining the genetic basis of a phenotype by leveraging genetic variation or divergence in a developmental, physiological, or behavioral trait. The work could also involve studying the evolutionary phenomena (e.g., reproductive isolation, adaptation, sexual dimorphism, behavior) that reveal an indirect link between genotype and a trait of interest. When the phenotype diverges across evolutionarily distinct lineages, this genotype-to-phenotype problem can be addressed using phylogenetic genotype-to-phenotype (PhyloG2P) mapping, which uses genetic signatures and convergent phenotypes on a phylogeny to infer the genetic bases of traits. The PhyloG2P approach has proven powerful in revealing key genetic changes associated with diverse traits, including the mammalian transition to marine environments and transitions between major mechanisms of photosynthesis. However, there are several intermediate traits layered in between genotype and the phenotype of interest, including but not limited to transcriptional profiles, chromatin states, protein abundances, structures, modifications, metabolites, and physiological parameters. Each intermediate trait is interesting and informative in its own right, but synthesis across data types has great promise for providing a deep, integrated, and predictive understanding of how genotypes drive phenotypic differences and convergence. We argue that an expanded PhyloG2P framework (the PhyloG2P matrix) that explicitly considers intermediate traits, and imputes those that are prohibitive to obtain, will allow a better mechanistic understanding of any trait of interest. Furthermore, this approach provides a proxy for functional validation and mechanistic understanding in organisms where laboratory manipulation is impractical.

59 BASIC BIOLOGICAL SCIENCES↗

Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search

Optimization problems become fundamentally challenging as the number of variables increases. Because the volume of the search space grows exponentially, classical algorithms frequently fail to locate the global minimum of non-convex functions. While quantum optimization offers a potential alternative, mapping continuous problems onto near-term quantum hardware introduces severe scaling limits and barren plateaus. To bridge this gap, we propose the Distributed Quantum-Enhanced Optimization (D-QEO) framework. Instead of forcing the quantum processor to find the exact minimum, we use it simply as a topographical preconditioner. The QPU maps the landscape to locate the most promising basin of attraction, generating high-quality seed points for a classical GPU-accelerated solver to refine. To make this approach viable for utility-scale problems, we exploit the mathematical structure of separable functions. This allows us to cut a 50-qubit (i.e., $2^{50}$) global search space into independent and manageable sub-spaces using 5-qubit subcircuits. By executing these fragments concurrently with CUDA-Q, we completely bypass the overhead of cross-register entanglement and classical tensor knitting for separable functions. Benchmarks on the 10-dimensional Rastrigin and Ackley functions show that D-QEO prevents the exponential failure rates observed in purely classical algorithms. Furthermore, this quantum warm-start significantly reduces the number of classical BFGS iterations required to converge, providing a highly practical blueprint for utilizing near-term quantum resources in complex global search.

Soos, Dominik [Old Dominion U.]↗

Massive νs through the CNN lens: interpreting the field-level neutrino mass information in weak lensing

Modern cosmological surveys probe the Universe deep into the nonlinear regime, where massive neutrinos suppress cosmic structure. Traditional cosmological analyses, which use the 2-point correlation function to extract information, are no longer optimal in the nonlinear regime, and there is thus much interest in extracting beyond-2-point information to improve constraints on neutrino mass. Quantifying and interpreting the beyond-2-point information is thus a pressing task. We study the field-level information in weak lensing convergence maps using convolution neural networks. We find that the network performance increases as higher source redshifts and smaller scales are considered — investigating up to a source redshift of 2.5 and ℓ max ≃ 10 4 — verifying that massive neutrinos leave a distinct effect on weak lensing. However, the performance of the network significantly drops after scaling out the 2-point information from the maps, implying that most of the field-level information can be found in the 2-point correlation function alone. We quantify these findings in terms of the likelihood ratio and also use Integrated Gradient saliency maps to interpret which parts of the map the network is learning the most from. We find that, in the absence of noise, the network extracts a similar amount of information from the most overdense and underdense regions. However, upon adding noise, the information in underdense regions is distorted as noise disproportionately washes out void-like structures.

Golshan, Malika [University of California, Berkele↗

On the Representativity of Electrode Microstructure Parameters and Their Electrochemical Response for Lithium Ion Batteries

Lithium-ion battery electrochemical models require an accurate description of the electrodes microstructures to be predictive that can be achieved through nanoscale imaging. Such observations are however limited by their field of view (FOV), as they provide only a subset of the whole electrode volume that does not necessarily represent the whole electrode microstructure heterogeneity, and therefore can bias the microstructure analysis. A microstructure scale electrochemical model was used to investigate lithium plating onset, material non-uniform utilization, and in-plane heterogeneities for an NMC-graphite full cell. To evaluate the representativeness, and thus relevance, of these model predictions, a coupled representativity analysis has been performed on the microstructure parameters and, in a novel way, on the full cell electrochemical response. Electrode microstructure parameters representativeness has been first quantified using the representative volume element (RVE) methodology. The RVE major flaw is that ultimately it can only conclude if a FOV contains representative subvolumes of the FOV, but not if the FOV itself is representative of the electrode volume. Analysis can conclude negatively ('FOV is not representative'), but not positively ('FOV is representative'). One major contribution of this work was to quantify the convergence of the RVE size with the FOV, to actually investigate the FOV representativeness and thus partly remedy this intrinsic limitation. The analysis determined that performing a standard RVE calculation, without exploring its FOV convergence, is likely to strongly underestimate the actual RVE size. The new RVE methodology has been automated in the NREL open-source Microstructure Analysis Toolbox (MATBOX) and is available to the battery community. Representativeness of microstructure parameters is however only an intermediate step, as the end-results of an electrochemical model are performances predictions. Indeed, what is the practical consequence of a given deviation for a microstructure parameter? The microstructure parameter deviation propagations to the 3D microstructure scale electrochemical response have been then quantified for different charge rates. This defines a threshold for the microstructure parameters FOV for a desired maximum deviation of the electrochemical response. Such deviation propagation analysis is analogous to error propagation analysis and is necessary to determine the relevance of microstructure scale model predictions for macroscale predictions. Electrochemical model shows cell representative section areas are increasing with C-rate, due to higher in-plane heterogeneities, indicating larger FOVs are required specifically for fast charge modeling. Therefore, we introduced the novel concept of electrochemical RVE (eRVE) that is a function of the operating conditions (thus defined as a dynamic RVE), with an increasing dependence with the C-rate. Representativity analysis of the investigated cell determined a FOV of 144.4 x 54.4 m2 is large enough to establish a convergence on the representative section areas for low to intermediate C-rate (=2.5C), but not large enough to conclude for higher rates. This work aims to emphasize the importance of representativity analysis for LIB electrode microstructures, as it is required to estimate the error, and thus the relevance, of microstructure parameters intended to be used in macroscale models. The methodology and results can help researchers to select the relevant imaging and associated FOV required to provide accurate enough microstructure parameters.

ADVANCED PROPULSION SYSTEMS↗

Diverse signatures of convergent evolution in cactus-associated yeasts

Many distantly related organisms have convergently evolved traits and lifestyles that enable them to live in similar ecological environments. However, the extent of phenotypic convergence evolving through the same or distinct genetic trajectories remains an open question. Here, we leverage a comprehensive dataset of genomic and phenotypic data from 1,049 yeast species in the subphylum Saccharomycotina (Kingdom Fungi, Phylum Ascomycota) to explore signatures of convergent evolution in cactophilic yeasts, ecological specialists associated with cacti. We inferred that the ecological association of yeasts with cacti arose independently approximately 17 times. Using a machine learning–based approach, we further found that cactophily can be predicted with 76% accuracy from both functional genomic and phenotypic data. The most informative feature for predicting cactophily was thermotolerance, which we found to be likely associated with altered evolutionary rates of genes impacting the cell envelope in several cactophilic lineages. We also identified horizontal gene transfer and duplication events of plant cell wall–degrading enzymes in distantly related cactophilic clades, suggesting that putatively adaptive traits evolved independently through disparate molecular mechanisms. Notably, we found that multiple cactophilic species and their close relatives have been reported as emerging human opportunistic pathogens, suggesting that the cactophilic lifestyle—and perhaps more generally lifestyles favoring thermotolerance—might preadapt yeasts to cause human disease. This work underscores the potential of a multifaceted approach involving high-throughput genomic and phenotypic data to shed light onto ecological adaptation and highlights how convergent evolution to wild environments could facilitate the transition to human pathogenicity.

59 BASIC BIOLOGICAL SCIENCES↗

Moving from Information Assurance to Functional Assurance with Engineered Controls

Cyber threats to operational technology demand more than traditional IT defenses—they require full-spectrum mission assurance. Cyber-Informed Engineering (CIE) is an approach that embeds engineered controls into system design to ensure critical functions remain safe and reliable, even under attack. Unlike conventional cybersecurity tools, engineered controls act directly on physical processes to prevent unacceptable outcomes such as equipment damage or mission failure. This session will outline the CIE framework and share examples of consequence-based design that deliver true resilience, not just fail-safe behaviors. Attendees will learn how to integrate these principles into the engineering lifecycle to support resilient-by-design architectures and inform emerging standards. This talk sets the stage for the panel discussion on advancing CIE across sectors as digital and physical systems converge.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

The eXtended virtual element method for elliptic problems with weakly singular solutions

This paper introduces a novel eXtended virtual element method, an extension of the conforming virtual element method. The X-VEM is formulated by incorporating appropriate enrichment functions in the local spaces. The method is designed to handle highly generic enrichment functions, including singularities arising from fractured domains. By achieving consistency on the enrichment space, the method is proven to achieve arbitrary approximation orders even in the presence of singular solutions. The paper includes a complete convergence analysis under general assumptions on mesh regularity, and numerical experiments validating the method’s accuracy on various mesh families, demonstrating optimal convergence rates in the L 2 - and H 1 - norms on fractured or L-shaped domains.

97 MATHEMATICS AND COMPUTING↗

Parallel derivative-free optimization for simulation-based design of behind-the-meter energy systems

In this work, the integrated design and dispatch of behind-the-meter or distributed resources (e.g. stationary battery storage and solar PV generation) is considered. A simulation-based framework is employed, generating high-fidelity results with closed-loop predictive control at a fine resolution, at the expense of high computational cost (several minutes to a few hours per design point). To address this challenge, parallel derivative-free design methods are considered. Four methods are compared, including state-of-the-art surrogate-based methods (Radial-Basis Functions and Gaussian processes) and sampling strategies, an evolutionary-based method, and a simple sequential grid refinement method. As a case study, two types of design problem with increasing complexity are considered, namely, the design of behind-the-meter resources (three design variables) and the inclusion of grid capacity (four design variables). The second yields a constrained design problem for which violations can only be determined after solving the computationally expensive simulation. For the three-dimensional case, all methods present a good performance, achieving a solution within 1% of the optimum after the first iteration, with the sequential grid refinement exhibiting the fastest convergence and achieving the best final objective value. This indicates that the parallel evaluation of multiple sampling points may be more important than the choice of method for small decision spaces. For the four-dimensional constrained case, the Genetic Algorithm presents the best tradeoff between performance and computational effort, while the rough objective function terrain generated by constraint violation penalties reduces the performance of surrogate-based methods. Contour plots with flat regions indicate flexibility in the optimal design and highlight the importance of characterizing the solution space.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Transverse momentum distributions at large- x

We investigate the collinear matching of transverse momentum dependent (TMD) distributions at large values of x, computing and resumming the leading large-x asymptotics for matching coefficients. The large-x resummation is done directly within TMD distributions, ensuring the process-independence of the result. The derived resummation formulas are valid for all TMD distributions (except the pretzelosity). Their application improves perturbative convergence, provides practical estimation for unknown higher-order contributions, and sets restrictions for the nonperturbative part of models. Using the known anomalous dimensions, resummation can reach N 3 LL, often exceeding the accuracy of known coefficient functions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry↗

Explainable physics-based constraints on reinforcement learning for accelerator optimization

We present a reinforcement learning (RL) framework for optimizing particle accelerator experiments that builds explainable physics-based constraints on agent behavior. The goal is to increase transparency and trust by letting users verify that the agent’s decision-making process incorporates suitable physics. Our algorithm uses a learnable surrogate function for physical observables, such as energy, and uses them to fine-tune how actions are chosen. This surrogate can be represented by a neural network or by an interpretable sparse dictionary model. We test our algorithm on a range of particle accelerator optimization environments designed to emulate the Continuous Electron Beam Accelerator Facility at Jefferson Lab. By examining the mathematical form of the learned constraint function, we are able to confirm the agent has learned to use the established physics of each environment. In addition, we find that the introduction of a physics-based surrogate enables our RL algorithms to reliably converge for difficult high-dimensional accelerator optimization environments.

explainability↗

Reducing measurement costs by recycling the Hessian in adaptive variational quantum algorithms

Abstract Adaptive protocols enable the construction of more efficient state preparation circuits in variational quantum algorithms (VQAs) by utilizing data obtained from the quantum processor during the execution of the algorithm. This idea originated with Adaptive Derivative-Assembled Problem-Tailored variational quantum eigensolver (ADAPT-VQE), an algorithm that iteratively grows the state preparation circuit operator by operator, with each new operator accompanied by a new variational parameter, and where all parameters acquired thus far are optimized in each iteration. In ADAPT-VQE and other adaptive VQAs that followed it, it has been shown that initializing parameters to their optimal values from the previous iteration speeds up convergence and avoids shallow local traps in the parameter landscape. However, no other data from the optimization performed at one iteration is carried over to the next. In this work, we propose an improved quasi-Newton optimization protocol specifically tailored to adaptive VQAs. The distinctive feature in our proposal is that approximate second derivatives of the cost function are recycled across iterations in addition to optimal parameter values. We implement a quasi-Newton optimizer where an approximation to the inverse Hessian matrix is continuously built and grown across the iterations of an adaptive VQA. The resulting algorithm has the flavor of a continuous optimization where the dimension of the search space is augmented when the gradient norm falls below a given threshold. We show that this inter-optimization exchange of second-order information leads the approximate Hessian in the state of the optimizer to be consistently closer to the exact Hessian. As a result, our method achieves a superlinear convergence rate even in situations where the typical implementation of a quasi-Newton optimizer converges only linearly. Our protocol decreases the measurement costs in implementing adaptive VQAs on quantum hardware as well as the runtime of their classical simulation.

Ramôa, Mafalda (ORCID:0000000302187801)↗

Large Angle Rocking Beam Electron Diffraction Utilizing Electron Direct Detector

Electron diffraction of a crystal is fundamentally a function of the potential of that crystal. The intensity of electron diffraction patterns is, as a result, sensitive to the charge density of atoms and bonding inside crystals. Experimentally, the traditional method to probe this information is quantitative Convergent Beam Electron Diffraction (CBED). Quantitative CBED or QCEBD is a method which uses dynamic diffraction theory to quantify CBED intensities and to extract information about crystal structure and bonding. The sensitivity of dynamical scattering is leveraged to measure crystal symmetry and crystal structure factors. At its limit, electron structure factors are measured at high accuracy allowing access to chemical bonding information, which corroborate theoretical calculations through multipole model refinements of the experimental charge density. The primary limit to QCBED is the requirement of a non-overlapping convergent beam. This subsequently limits QCBED to crystals with small unit cells which are stable under the focused probe.

Busch, Robert↗

First-principles study of the Stark shift effect on the zero-phonon line of the NV center in diamond

Point defects in semiconductors are attractive candidates for quantum information science applications owing to their ability to act as spin-photon interface or single-photon emitters. However, the coupling between the change of dipole moment upon electronic excitation and stray electric fields in the vicinity of the defect, an effect known as Stark shift, can cause significant spectral diffusion in the emitted photons. In this work, using first principles computations, we revisit the methodology to compute the Stark shift of point defects up to the second order. The approach consists of applying an electric field on a defect in a slab and monitoring the changes in the computed zero-phonon line (i.e., difference in energy between the ground and excited state) obtained from constraining the orbital occupations (constrained-DFT). Here, we study the Stark shift of the negatively charged nitrogen-vacancy (NV) center in diamond using this slab approach. We discuss and compare two approaches to ensure a negatively charged defect in a slab and we show that converged values of the Stark shift measured by the change in dipole moment between the ground and excited states (Δ⁢μ) can be obtained. We obtain a Stark shift of Δ⁢μ = 2.68⁢D using the semilocal GGA-PBE functional and of Δ⁢μ = 2.23⁢D using the HSE hybrid functional. These values are in good agreement with experimental results. We also show that modern theory of polarization can be used on constrained-DFT to obtain Stark shifts in very good agreement with the slab computations.

36 MATERIALS SCIENCE↗

Datasets for Custom-trained Machine-learning Interatomic Potentials: Nitric Acid Aqueous Solution

This dataset was generated using an iterative active learning strategy with the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials (MLIPs) for aqueous nitric acid. Each active-learning cycle consisted of three stages: (1) training, (2) exploration, and (3) labeling. The initial training set comprised approximately 800 randomly selected configurations from a previous study by Lewis et al. (https://doi.org/10.1021/jp205510q), which investigated nitric acid solutions at 2, 3, 4, and 5 mol/L. For all configurations, single-point calculations of atomic forces and total energies were performed at the quantum density functional theory BLYP-D2 and PBE-D3 levels of theory using the CP2K Quickstep module. Valence electrons were treated explicitly, while core electrons on all atoms were represented by norm-conserving Goedecker–Teter–Hutter (GTH) pseudopotentials. Long-range dispersion interactions were accounted for using Grimme dispersion corrections. Wave functions were expanded in a mixed Gaussian-and-plane-wave scheme using TZV2P-MOLOPT basis sets for all elements and an 800 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent field convergence was accelerated using orbital transformation and Direct Inversion in the Iterative Subspace, with a convergence threshold of 10^{-6}. All single-point calculations were carried out in periodic orthorhombic cells whose dimensions match those of the molecular configurations sampled from earlier trajectories. The CELL_REF keyword in CP2K was used to define a fixed reference cell, ensuring consistency in the reference data used for MLIP training, particularly when cell fluctuations are present in NpT simulations. The resulting high-fidelity energies and forces constitute the ground-truth labels used to train the MLIPs contained in this dataset.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

Custom-trained Machine-learning Interatomic Potentials: ZnCl2 Aqueous Solution

This dataset was generated using an iterative active-learning strategy implemented in the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials for aqueous ZnCl2 solutions. Each active-learning cycle consisted of three stages: training, exploration, and labeling. The initial training set combined configurations generated in this work from enhanced-sampling ab initio molecular dynamics simulations with configurations from a previously reported neural-network-potential study of aqueous ZnCl2. The enhanced-sampling ab initio molecular dynamics simulations involved Zn–Cl separation and the chloride coordination number around Zn²? as collective variables. These configurations served as the seed dataset. Subsequent active-learning cycles expanded the training set by identifying and labeling configurations that were poorly represented by the current models, thereby improving coverage of ion-association states and changes in local coordination and charge-state environments relevant to the solution free-energy landscape. For all selected configurations, single-point calculations of the total energies and atomic forces were performed within density functional theory using the CP2K Quickstep module. Reference calculations employed the revPBE-D3 and r2SCAN exchange-correlation functionals. Motivated by recent work on aqueous Zn²?, the main revPBE calculations omitted D3 dispersion contributions involving Zn²?, while retaining the D3 correction for water and chloride. For comparison, fully dispersion-corrected revPBE-D3 reference calculations were also performed, with D3 applied to all species, including Zn²?. Valence electrons were treated explicitly, while core electrons were represented using norm-conserving Goedecker–Teter–Hutter pseudopotentials. The wave functions were expanded using the mixed Gaussian-and-plane-wave scheme with TZV2P-MOLOPT basis sets for all elements and a 600 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent-field convergence was accelerated using the orbital-transformation and Direct Inversion in the Iterative Subspace algorithms, with a convergence threshold of 10?6. All single-point calculations were performed in periodic orthorhombic cells. The CELL_REF keyword in CP2K was used to define a fixed reference cell with a box length of 25 Å. This treatment ensured a consistent reference for configurations extracted from NpT trajectories with fluctuating cell dimensions. The resulting DFT energies and atomic forces constitute the ground-truth labels used to train the MLIPs. The resulting MLIP was trained for aqueous ZnCl2 solutions spanning concentrations from 0 to 30 molal and a broad pH range, from strongly acidic to strongly basic conditions. Representative examples of configurations included in the MLIP training dataset are provided below. These include 1) Representative configurations from the dataset labeled at the revPBE-D3 level, with D3 dispersion interactions involving Zn2+ excluded (revPBE-wo-D3). 2) Representative configurations from the dataset labeled at the fully dispersion-corrected revPBE-D3 level, with D3 interactions applied to all species, including Zn2+ (revPBE-D3). 3) Representative configurations from the dataset labeled at the r2SCAN level of theory (r2SCAN).

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗