Search NASA⌕ Search

SEARCH · Search NASA

Results for “relevancy algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Linear Subpixel Learning Algorithm for Land Cover Classification from WELD using High Performance Computing

In this work, we use a Fully Constrained Least Squares Subpixel Learning Algorithm to unmix global WELD (Web Enabled Landsat Data) to obtain fractions or abundances of substrate (S), vegetation (V) and dark objects (D) classes. Because of the sheer nature of data and compute needs, we leveraged the NASA Earth Exchange (NEX) high performance computing architecture to optimize and scale our algorithm for large-scale processing. Subsequently, the S-V-D abundance maps were characterized into 4 classes namely, forest, farmland, water and urban areas (with NPP-VIIRS-national polar orbiting partnership visible infrared imaging radiometer suite nighttime lights data) over California, USA using Random Forest classifier. Validation of these land cover maps with NLCD (National Land Cover Database) 2011 products and NAFD (North American Forest Dynamics) static forest cover maps showed that an overall classification accuracy of over 91 percent was achieved, which is a 6 percent improvement in unmixing based classification relative to per-pixel-based classification. As such, abundance maps continue to offer an useful alternative to high-spatial resolution data derived classification maps for forest inventory analysis, multi-class mapping for eco-climatic models and applications, fast multi-temporal trend analysis and for societal and policy-relevant applications needed at the watershed scale.

Subpixel↗

Visual and Inertial Datasets for an eVTOL Aircraft Approach and Landing Scenario

A National Aeronautics and Space Administration (NASA) project developing computer vision algorithms for autonomous flight is producing real-world datasets with cameras mounted on aircraft. In related domains, such as autonomous driving, open datasets are key to innovation and advancement in computer vision and autonomous perception for future Advanced Air Mobility (AAM) operations. Few vision datasets, however, are publicly available in the aviation context. This paper introduces preliminary datasets containing several examples of approach and landing scenarios. The platform aircraft include a multirotor small unmanned aerial system (sUAS) and a crewed helicopter as surrogates for future electric vertical take-off and landing (eVTOL) aircraft. The dataset provides video imagery with associated inertial navigation system-global positioning system (INS-GPS) position and attitude estimates and other sensors. Surveyed locations of the visual features of the landing area are included. This dataset is the first to be released in an ongoing effort to collect and share large, diverse datasets relevant to autonomous aviation; community critique that can inform and improve future flight campaigns is welcome.

Nelson Brown↗

Detailed Analysis of Indian Summer Monsoon Rainfall Processes with Modern/High-Quality Satellite Observations

We examine, in detail, Indian Summer Monsoon rainfall processes using modernhigh quality satellite precipitation measurements. The focus here is on measurements derived from three NASA cloud and precipitation satellite missionslinstruments (TRMM/PR&TMI, AQUNAMSRE, and CLOUDSATICPR), and a fourth TRMM Project-generated multi-satellite precipitation measurement dataset (viz., TRMM standard algorithm 3b42) -- all from a period beginning in 1998 up to the present. It is emphasized that the 3b42 algorithm blends passive microwave (PMW) radiometer-based precipitation estimates from LEO satellites with infi-ared (IR) precipitation estimates from a world network of CEO satellites (representing -15% of the complete space-time coverage) All of these observations are first cross-calibrated to precipitation estimates taken from standard TRMM combined PR-TMI algorithm 2b31, and second adjusted at the large scale based on monthly-averaged rain-gage measurements. The blended approach takes advantage of direct estimates of precipitation from the PMW radiometerequipped LEO satellites -- but which suffer fi-om sampling limitations -- in combination with less accurate IR estimates from the optical-infrared imaging cameras on GEO satellites -- but which provide continuous diurnal sampling. The advantages of the current technologies are evident in the continuity and coverage properties inherent to the resultant precipitation datasets that have been an outgrowth of these stable measuring and retrieval technologies. There is a wealth of information contained in the current satellite measurements of precipitation regarding the salient precipitation properties of the Indian Summer Monsoon. Using different datasets obtained from the measuring systems noted above, we have analyzed the observations cast in the form of: (1) spatially distributed means and variances over the hierarchy of relevant time scales (hourly I diurnally, daily, monthly, seasonally I intra-seasonally, and inter-annually), (2) time series at these different time scales taken as area-averages over the hierarchy of relevant space scales (Indian sub-Division, Indian sub-continent, and Circumambient Indian Ocean), (3) principal autocorrelation and cross-correlation structures over various monsoon space-time domains, (4) diurnally modulated amplitude-phase properties of rain rates over different monsoon space-time domains, (5) foremost rain rate probability distributions intrinsic to monsoon precipitation, and (6) behavior of extreme events including occurrences of flood and drought episodes throughout the course of inter-annual monsoon processes.

Smith, Eric A.↗

Algorithm Theoretical Basis Document (ATBD) - Stream Stage Measurements: V2.5.1 Water Level Products from Satellite Radar Altimetry

In response to the 2018 NASA ROSES Applied Sciences/Water Resources (NASA HQ Program Official: Dr. Brad Doorn) call for proposals, the “Integration of Remotely Sensed Streamflow Data into Alaska Water Resource Management Agency Operations” project with Principal Investigator (PI) Jack Eggleston USGS, was successful, and had the ultimate goal of creating a series of remotely sensed or derived Alaska river parameters for integration into NWIS. These parameters included surface water height and average reach surface water slope (from altimetry), average reach width (from Landsat imagery), and an associated river discharge derived via theoretical means. The surface water height products were required to have both archival and near real time components, noting the availability of ~25years of potential measurements, and accepting the temporal resolution (10-35days) of the suite of radar altimeters. Each surface water level product was expected to be a continuous time series of observation with a sufficient accuracy to highlight monthly, seasonal and interannual variation. The designated set of river reaches were chosen for their geographical distribution, their reach width, and the presence of a radar altimeter mission satellite overpass. This document describes the procedure associated with the creation of these altimetric surface water level products and is relevant to product Version 2.5.1 available from the Global Water Monitor (GWM) web portal.

Altimetry↗

FORESTR: Finding, Organizing, Representing, Explaining, Summarizing, and Thinning Random forests

Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.

97 MATHEMATICS AND COMPUTING↗

New Ways of Facilitating Improved Data Discovery and Access for NASA's Suborbital Earth Science Observations

NASA conducts field research in various Earth Science disciplines utilizing airborne and other non-satellite platforms to acquire in situ and remotely sensed observations indicative of physical processes across a range of scales. Field efforts are key in the development and validation of instruments and satellite algorithm refinements. The heterogeneous data, with a range of file formats, scales, and acquisition methods, support research in several science areas. NASA’s archive process assigns data products to discipline-oriented Distributed Active Archive Centers (DAACs) for stewardship. Over time, individual DAACs have developed tools for data browsing and serving disparate user bases. As science becomes more interdisciplinary, researchers need to incorporate observations from multiple campaigns, and multiple DAACs, into their work. Motivated in part by this shifting paradigm of needs, the Catalog of Archived Suborbital Earth Science Investigations (CASEI) was created. CASEI provides a single starting point to browse, search, and discover airborne and field data. Contextual metadata are organized and inter-linked allowing intuitive, integrated exploration across all NASA DAACs. Campaign science objectives, platform and instrument configurations, geographical details, geophysical concepts, and more are tracked in CASEI’s database, facilitating multi-parameter search, browse, and discovery of relevant data products. Researchers are able to directly access associated data products, via DOI links, regardless of the DAAC where they reside. Significant events, key time periods of high science interest within the longer-duration campaign effort, are also indicated and allow for a more efficient identification of critical data subsets. This presentation describes CASEI’s development, intensive metadata curation process, and demonstrates the web interface experience. Initial content metrics and plans for continued maintenance will also be discussed.

metadata↗

ELISA: A Tool for Optimization of Rotor Hover Performance at Low Reynolds Number in the Mars Atmosphere

The Evolutionary aLgorithm for Iterative Studies of Aeromechanics (ELISA) was developed in support of the Rotorcraft Optimization for the Advancement of Mars eXploration (ROAMX) project. ELISA was developed to enable aerodynamic rotor hover optimization for low Reynolds number flows in the Mars atmosphere. The first objective of the algorithm allows for unconventional airfoil parameterization and multi-objective airfoil geometry optimization using OVERFLOW. The Pareto-optimal airfoil sets are converted to a set of Pareto-optimal airfoil decks, providing the lowest drag air foil geometry for each angle of attack, removing the need to arbitrarily select the airfoils to be used in the rotor optimization. The second objective allows for rotor geometry optimization with simultaneous maximization of blade loading and minimization of rotor power using the comprehensive analysis CAMRADII. The result is a Pareto-optimal rotor set, providing the lowest power rotor for each attainable blade loading, and one of the first tools for hover-optimized rotors for high-subsonic low Reynolds number conditions. The airfoil thickness can be modified after the airfoil optimization is complete, allowing for a post-airfoil-optimization adjustment of blade thickness to facilitate conforming to external structural analyses requirements. The relevance of the code is demonstrated with case studies for the ROAMX rotor optimization for Ingenuity-sized single rotors in the Mars atmosphere, a performance study optimizing the chord and twist of Ingenuity’s coaxial rotor resulting in the Sample Recovery Helicopters candidate rotor, and high-subsonic low Reynolds number airfoil optimization providing novel insights for higher-efficiency low Reynolds number airfoil geometries and flow physics.

ELISA↗

Adaptive Numerical Dissipation Control in High Order Schemes for Multi-D Non-Ideal MHD

The required type and amount of numerical dissipation/filter to accurately resolve all relevant multiscales of complex MHD unsteady high-speed shock/shear/turbulence/combustion problems are not only physical problem dependent, but also vary from one flow region to another. In addition, proper and efficient control of the divergence of the magnetic field (Div(B)) numerical error for high order shock-capturing methods poses extra requirements for the considered type of CPU intensive computations. The goal is to extend our adaptive numerical dissipation control in high order filter schemes and our new divergence-free methods for ideal MHD to non-ideal MHD that include viscosity and resistivity. The key idea consists of automatic detection of different flow features as distinct sensors to signal the appropriate type and amount of numerical dissipation/filter where needed and leave the rest of the region free from numerical dissipation contamination. These scheme-independent detectors are capable of distinguishing shocks/shears, flame sheets, turbulent fluctuations and spurious high-frequency oscillations. The detection algorithm is based on an artificial compression method (ACM) (for shocks/shears), and redundant multiresolution wavelets (WAV) (for the above types of flow feature). These filters also provide a natural and efficient way for the minimization of Div(B) numerical error.

Yee, H. C.↗

Towards large-scale quantum optimization solvers with few qubits

Quantum computers hold the promise of more efficient combinatorial optimization solvers, which could be game-changing for a broad range of applications. However, a bottleneck for materializing such advantages is that, in order to challenge classical algorithms in practice, mainstream approaches require a number of qubits prohibitively large for near-term hardware. Here we introduce a variational solver for MaxCut problems over $m={{\mathcal{O}}}({n}^{k})$ binary variables using only n qubits, with tunable k > 1. The number of parameters and circuit depth display mild linear and sublinear scalings in m , respectively. Moreover, we analytically prove that the specific qubit-efficient encoding brings in a super-polynomial mitigation of barren plateaus as a built-in feature. Altogether, this leads to high quantum-solver performances. For instance, for m = 7000, numerical simulations produce solutions competitive in quality with state-of-the-art classical solvers. In turn, for m = 2000, experiments with n = 17 trapped-ion qubits feature MaxCut approximation ratios estimated to be beyond the hardness threshold 0.941. Our findings offer an interesting heuristics for quantum-inspired solvers as well as a promising route towards solving commercially-relevant problems on near-term quantum devices.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Tactical Analysis for Calculating Contextual Risk at Boundaries: Summary of Laboratory Directed Research & Development Effort

The Tactical Analysis for Calculating Contextual Risk at Boundaries (TACCRAB) tool is an innovative digital twin (DT) platform and automated risk algorithm designed to transform operational decision-making in structured screening environments, with an initial focus on Southern Border Land Ports of Entry (POEs). The invention provides integration points for advanced artificial intelligence, predictive modeling, and real-time data analysis to produce a comprehensive risk management tool that enables proactive, data-informed security strategies. The core inventive features of TACCRAB center on its unique risk algorithm, which dynamically calculates contextual risk by synthesizing historical data, near real-time streaming data from the checkpoints themselves, and AI-generated predictions. Unlike traditional risk assessment methods, TACCRAB utilizes a DT to provide comprehensive operational insights, allowing stakeholders to visualize, simulate, and optimize checkpoint configurations with unprecedented speed and contextual awareness. TACCRAB's key innovation lies in its ability to combine multiple complex inputs - including technology detection probabilities, resource availability, screening pathway characteristics, and threat actor behavioral patterns - into a unified risk calculation and update these inputs based on changing operational and environmental conditions. By leveraging a DT that continuously updates and learns from linked data, TACCRAB can suggest adaptive mitigation strategies that minimize risk while maintaining operational efficiency. Particularly novel is the platform's approach to decision support, which goes beyond static risk assessment. The DT provides dynamic metrics such as wait times, resource allocation effectiveness, and potential emerging threat scenarios, enabling users to view sophisticated, relevant what-if simulations and optimize checkpoint operations in near real-time. The system's architecture allows for generalized application across different screening environments, such as secure facilities, ports of entry, and soft targets, making it a versatile tool for security and operational management. The invention distinguishes itself through its comprehensive integration of predictive modeling, AI-driven pattern discovery, and user-friendly interface design. By combining these elements, TACCRAB transforms complex risk data into actionable insights, supporting decision-makers at various organizational levels - from booth agents making split-second screening decisions to checkpoint managers optimizing the day's resource allocation to strategic planners managing long-term investments.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Modeling Guidelines for Code Generation in the Railway Signaling Context

Modeling guidelines constitute one of the fundamental cornerstones for Model Based Development. Their relevance is essential when dealing with code generation in the safety-critical domain. This article presents the experience of a railway signaling systems manufacturer on this issue. Introduction of Model-Based Development (MBD) and code generation in the industrial safety-critical sector created a crucial paradigm shift in the development process of dependable systems. While traditional software development focuses on the code, with MBD practices the focus shifts to model abstractions. The change has fundamental implications for safety-critical systems, which still need to guarantee a high degree of confidence also at code level. Usage of the Simulink/Stateflow platform for modeling, which is a de facto standard in control software development, does not ensure by itself production of high-quality dependable code. This issue has been addressed by companies through the definition of modeling rules imposing restrictions on the usage of design tools components, in order to enable production of qualified code. The MAAB Control Algorithm Modeling Guidelines (MathWorks Automotive Advisory Board)[3] is a well established set of publicly available rules for modeling with Simulink/Stateflow. This set of recommendations has been developed by a group of OEMs and suppliers of the automotive sector with the objective of enforcing and easing the usage of the MathWorks tools within the automotive industry. The guidelines have been published in 2001 and afterwords revisited in 2007 in order to integrate some additional rules developed by the Japanese division of MAAB [5]. The scope of the current edition of the guidelines ranges from model maintainability and readability to code generation issues. The rules are conceived as a reference baseline and therefore they need to be tailored to comply with the characteristics of each industrial context. Customization of these recommendations has been performed for the automotive control systems domain in order to enforce code generation [7]. The MAAB guidelines have been found profitable also in the aerospace/avionics sector [1] and they have been adopted by the MathWorks Aerospace Leadership Council (MALC). General Electric Transportation Systems (GETS) is a well known railway signaling systems manufacturer leading in Automatic Train Protection (ATP) systems technology. Inside an effort of adopting formal methods within its own development process, GETS decided to introduce system modeling by means of the MathWorks tools [2], and in 2008 chose to move to code generation. This article reports the experience performed by GETS in developing its own modeling standard through customizing the MAAB rules for the railway signaling domain and shows the result of this experience with a successful product development story.

Ferrari, Alessio↗

New NDA Methods for Thorium Fuel Cycle Safeguards (Final Report)

This project developed portable Neutron Resonance Transmission Analysis (pNRTA) as a new non-destructive assay (NDA) method for thorium fuel cycles safeguards and other applications where multiple isotopes must be measured when present together. pNRTA leverages epithermal neutron resonances to assay multiple safeguards-relevant isotopes (e.g., 233 U and 235 U) when they are present together in a sample. Existing techniques are challenged by this task, driving the need for new active interrogation methods. With selected detectors, pNRTA works in high gamma-ray backgrounds from fission and activation products and 232 U progeny expected in thorium fuel cycle samples. This project leveraged a pNRTA system developed at Pacific Northwest National Laboratory (PNNL) and collaboration with the Massachusetts Institute of Technology (MIT). The system uses a commercially available deuterium-tritium (DT) neutron generator at short standoff (2 m). Key achievements in this project included: first-of-a-kind pNRTA quantitative measurements of 233 U oxide samples, an assessment of neutron detector technologies suitable for pNRTA in high gamma-ray background environments, experimentally demonstrating quantitative assay of samples containing 233 U and 235 U, and modeling studies showing the applicability of pNRTA to a wide range of material forms. Further, a custom algorithm was developed at MIT, which provided mean bias of 9% and relative standard deviation of 36% in assaying 233 U, 235 U, 238 U, and 232 Th content in eight measured samples. These outcomes form a solid technical basis for pNRTA as a new promising capability for international safeguards verification that is portable, non-destructive, quantitative, and isotopic specific.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Discretization and Preconditioning Algorithms for the Euler and Navier-Stokes Equations on Unstructured Meshes

Several stabilized demoralization procedures for conservation law equations on triangulated domains will be considered. Specifically, numerical schemes based on upwind finite volume, fluctuation splitting, Galerkin least-squares, and space discontinuous Galerkin demoralization will be considered in detail. A standard energy analysis for several of these methods will be given via entropy symmetrization. Next, we will present some relatively new theoretical results concerning congruence relationships for left or right symmetrized equations. These results suggest new variants of existing FV, DG, GLS, and FS methods which are computationally more efficient while retaining the pleasant theoretical properties achieved by entropy symmetrization. In addition, the task of Jacobean linearization of these schemes for use in Newton's method is greatly simplified owing to exploitation of exact symmetries which exist in the system. The FV, FS and DG schemes also permit discrete maximum principle analysis and enforcement which greatly adds to the robustness of the methods. Discrete maximum principle theory will be presented for general finite volume approximations on unstructured meshes. Next, we consider embedding these nonlinear space discretizations into exact and inexact Newton solvers which are preconditioned using a nonoverlapping (Schur complement) domain decomposition technique. Elements of nonoverlapping domain decomposition for elliptic problems will be reviewed followed by the present extension to hyperbolic and elliptic-hyperbolic problems. Other issues of practical relevance such the meshing of geometries, code implementation, turbulence modeling, global convergence, etc, will. be addressed as needed.

Barth, Timothy J.↗

Mobile LiDAR as a Tool for Terrestrial and Planetary Cave Exploration and Mapping

KNaCK (Kinematic Navigation and Cartography Knapsack) is a backpack-mounted mobile mapping system. It can map its surroundings in 3 dimensions and localize itself in space using a LiDAR (Light Detection and Ranging) sensor and SLAM (Simultaneous Localization and Mapping) algorithm. The KNaCK team is leveraging caves as a proving ground to refine technology for mapping and navigation on other worlds while simultaneously advancing the State of the Art for terrestrial cave exploration and study.

LiDAR↗

Bridging Control and Deployment: A Cross-Layer Analysis of Scalable Building Cluster Control

Building cluster control has emerged as a promising approach for enabling flexible and coordinated operation of distributed building systems, yet its transition from pilot demonstrations to routine grid-interactive operation remains limited. This paper argues that this gap cannot be explained by control algorithms alone. Instead, it arises from interacting barriers in communication infrastructure, data and semantic interoperability, uncertainty management, stakeholder participation, market design, and policy support. Accordingly, the paper reviews both technical and non-technical barriers to building cluster control. Technical challenges include heterogeneous devices and protocols, communication latency and reliability, distributed decision-making, and uncertainty propagation across aggregated loads. Non-technical barriers include user participation, stakeholder coordination, incentive allocation, and data governance. Existing solution approaches are synthesized, including semantic interoperability frameworks, edge and hierarchical communication architectures, distributed and transactive control strategies, uncertainty-aware optimization, policy mechanisms, and market reforms. Based on this analysis, two research directions are identified: testing infrastructures that can evaluate control performance under realistic multi-building conditions, and abstraction methods that allow building clusters to interact with other energy sectors through standardized flexibility representations. Overall, the paper provides a structured review of how building cluster control can move from isolated demonstrations toward reproducible, market-compatible, and grid-relevant implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Static Subspace Approximation for Random Phase Approximation Correlation Energies: Implementation and Performance

Developing theoretical understanding of complex reactions and processes at interfaces requires using methods that go beyond semilocal density functional theory to accurately describe the interactions between solvent, reactants and substrates. Methods based on many-body perturbation theory, such as the random phase approximation (RPA), have previously been limited due to their computational complexity. However, this is now a surmountable barrier due to the advances in computational power available, in particular through modern GPU-based supercomputers. In this work, we describe the implementation of RPA calculations within BerkeleyGW and show its favorable computational performance on large complex systems relevant for catalysis and electrochemistry applications. Our implementation builds off of the static subspace approximation which, by employing a compressed representation of the frequency dependent polarizability, enables the evaluation of the RPA correlation energy with significant acceleration and systematically controllable accuracy. We find that the computational cost of calculating the RPA correlation energy scales only linearly with system size for systems containing up to 50 thousand bands, and is expected to scale quadratically thereafter. We also show excellent strong scaling results across several supercomputers, demonstrating the performance and portability of this implementation.

algorithmic development↗

Discretization and Preconditioning Algorithms for the Euler and Navier-Stokes Equations on Unstructured Meshes

Several stabilized discretization procedures for conservation law equations on triangulated domains will be considered. Specifically, numerical schemes based on upwind finite volume, fluctuation splitting, Galerkin least-squares, and space discontinuous Galerkin discretization will be considered in detail. A standard energy analysis for several of these methods will be given via entropy symmetrization. Next, we will present some relatively new theoretical results concerning congruence relationships for left or right symmetrized equations. These results suggest new variants of existing FV, DG, GLS and FS methods which are computationally more efficient while retaining the pleasant theoretical properties achieved by entropy symmetrization. In addition, the task of Jacobian linearization of these schemes for use in Newton's method is greatly simplified owing to exploitation of exact symmetries which exist in the system. These variants have been implemented in the "ELF" library for which example calculations will be shown. The FV, FS and DG schemes also permit discrete maximum principle analysis and enforcement which greatly adds to the robustness of the methods. Some prevalent limiting strategies will be reviewed. Next, we consider embedding these nonlinear space discretizations into exact and inexact Newton solvers which are preconditioned using a nonoverlapping (Schur complement) domain decomposition technique. Elements of nonoverlapping domain decomposition for elliptic problems will be reviewed followed by the present extension to hyperbolic and elliptic-hyperbolic problems. Other issues of practical relevance such the meshing of geometries, code implementation, turbulence modeling, global convergence, etc. will be addressed as needed.

Barth, Timothy↗

Custom-trained Machine-learning Interatomic Potentials: ZnCl2 Aqueous Solution

This dataset was generated using an iterative active-learning strategy implemented in the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials for aqueous ZnCl2 solutions. Each active-learning cycle consisted of three stages: training, exploration, and labeling. The initial training set combined configurations generated in this work from enhanced-sampling ab initio molecular dynamics simulations with configurations from a previously reported neural-network-potential study of aqueous ZnCl2. The enhanced-sampling ab initio molecular dynamics simulations involved Zn–Cl separation and the chloride coordination number around Zn²? as collective variables. These configurations served as the seed dataset. Subsequent active-learning cycles expanded the training set by identifying and labeling configurations that were poorly represented by the current models, thereby improving coverage of ion-association states and changes in local coordination and charge-state environments relevant to the solution free-energy landscape. For all selected configurations, single-point calculations of the total energies and atomic forces were performed within density functional theory using the CP2K Quickstep module. Reference calculations employed the revPBE-D3 and r2SCAN exchange-correlation functionals. Motivated by recent work on aqueous Zn²?, the main revPBE calculations omitted D3 dispersion contributions involving Zn²?, while retaining the D3 correction for water and chloride. For comparison, fully dispersion-corrected revPBE-D3 reference calculations were also performed, with D3 applied to all species, including Zn²?. Valence electrons were treated explicitly, while core electrons were represented using norm-conserving Goedecker–Teter–Hutter pseudopotentials. The wave functions were expanded using the mixed Gaussian-and-plane-wave scheme with TZV2P-MOLOPT basis sets for all elements and a 600 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent-field convergence was accelerated using the orbital-transformation and Direct Inversion in the Iterative Subspace algorithms, with a convergence threshold of 10?6. All single-point calculations were performed in periodic orthorhombic cells. The CELL_REF keyword in CP2K was used to define a fixed reference cell with a box length of 25 Å. This treatment ensured a consistent reference for configurations extracted from NpT trajectories with fluctuating cell dimensions. The resulting DFT energies and atomic forces constitute the ground-truth labels used to train the MLIPs. The resulting MLIP was trained for aqueous ZnCl2 solutions spanning concentrations from 0 to 30 molal and a broad pH range, from strongly acidic to strongly basic conditions. Representative examples of configurations included in the MLIP training dataset are provided below. These include 1) Representative configurations from the dataset labeled at the revPBE-D3 level, with D3 dispersion interactions involving Zn2+ excluded (revPBE-wo-D3). 2) Representative configurations from the dataset labeled at the fully dispersion-corrected revPBE-D3 level, with D3 interactions applied to all species, including Zn2+ (revPBE-D3). 3) Representative configurations from the dataset labeled at the r2SCAN level of theory (r2SCAN).

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗