Search NASASearch

SEARCH · Search NASA

Results for “dimensionality reduction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Layered and Low-Dimensional Lead, Silver, and Bismuth Halide Perovskites Directed by Halogen-Substituted Spacer Cations

Hybrid organic–inorganic metal halides provide a diverse parameter space in which the optoelectronic properties can be tuned through the composition. The compositional tunability extends to the metal site, which can be expanded from single valent metals (e.g., Pb 2+ ) to multivalent metals (e.g., Ag + and Bi 3+ ), and the dimension (2D, 1D, or 0D). However, a deeper understanding of how the organic cations template these metal halide structures is needed. Here, we synthesize and study the structures of a series of new layered and low-dimensional metal (Pb, Ag, and Bi) halides templated by the halogenated aryl spacer cations 2-chlorobenzylammonium (2ClBZ) and 3-chloro-2-fluorobenzylammonium (3Cl2FBZ). We report new lead perovskites, (3Cl2FBZ) 2 PbBr 4 , (2ClBZ) 3 PbI 5 , and (3Cl2FBZ) 2 PbI 4 , and compare them to their silver and/or bismuth analogs (2ClBZ) 4 AgBiBr 8 , (3Cl2FBZ) 4 AgBiBr 8 , (2ClBZ) 3 Bi 2 I 9 , and (3Cl2FBZ) 4 Bi 2 I 10 . In all structures, the halogen-substituted cations result in 2D or “pseudo-2D” layering, but the different halogen substituents introduce different distortions (tilting, octahedral distortion) and dimensional reduction to 1D or 0D depending on the metal and halide compositions. Optical absorption measurements reveal the bandgaps are tunable through metal sites, dimension, and cations to different extents. Furthermore, the 1D (3Cl2FBZ) 4 Bi 2 I 10 crystallizes in the noncentrosymmetric space group Cmc2 1 and exhibits second-harmonic generation (SHG). Furthermore, the organic–inorganic interactions and resultant structural distortions examined here provide insights toward the engineering of noncentrosymmetry and dimensional control in hybrid metal halide perovskites.

Cations

Neural Network Machine Learning and Dimension Reduction for Data Visualization

Neural network machine learning in computer science is a continuously developing field of study. Although neural network models have been developed which can accurately predict a numeric value or nominal classification, a general purpose method for constructing neural network architecture has yet to be developed. Computer scientists are often forced to rely on a trial-and-error process of developing and improving accurate neural network models. In many cases, models are constructed from a large number of input parameters. Understanding which input parameters have the greatest impact on the prediction of the model is often difficult to surmise, especially when the number of input variables is very high. This challenge is often labeled the "curse of dimensionality" in scientific fields. However, techniques exist for reducing the dimensionality of problems to just two dimensions. Once a problem's dimensions have been mapped to two dimensions, it can be easily plotted and understood by humans. The ability to visualize a multi-dimensional dataset can provide a means of identifying which input variables have the highest effect on determining a nominal or numeric output. Identifying these variables can provide a better means of training neural network models; models can be more easily and quickly trained using only input variables which appear to affect the outcome variable. The purpose of this project is to explore varying means of training neural networks and to utilize dimensional reduction for visualizing and understanding complex datasets.

Liles, Charles A.

Co-orchestration of multiple instruments to uncover structure–property relationships in combinatorial libraries

The rapid growth of automated and autonomous instrumentation brings forth opportunities for the co-orchestration of multimodal tools that are equipped with multiple sequential detection methods or several characterization techniques to explore identical samples. This is exemplified by combinatorial libraries that can be explored in multiple locations via multiple tools simultaneously or downstream characterization in automated synthesis systems. In co-orchestration approaches, information gained in one modality should accelerate the discovery of other modalities. Correspondingly, an orchestrating agent should select the measurement modality based on the anticipated knowledge gain and measurement cost. Herein, we propose and implement a co-orchestration approach for conducting measurements with complex observables, such as spectra or images. The method relies on combining dimensionality reduction by variational autoencoders with representation learning for control over the latent space structure and integration into an iterative workflow via multi-task Gaussian Processes (GPs). This approach further allows for the native incorporation of the system's physics via a probabilistic model as a mean function of the GPs. We illustrate this method for different modes of piezoresponse force microscopy and micro-Raman spectroscopy on a combinatorial Sm-BiFeO3 library. However, the proposed framework is general and can be extended to multiple measurement modalities and arbitrary dimensionality of the measured signals.

47 OTHER INSTRUMENTATION

Machine-learned closure of URANS for stably stratified turbulence: connecting physical timescales & data hyperparameters of deep time-series models

Stably stratified turbulence (SST), a model that is representative of the turbulence found in the oceans and atmosphere, is strongly affected by fine balances between forces and becomes more anisotropic in time for decaying scenarios. Moreover, there is a limited understanding of the physical phenomena described by some of the terms in the Unsteady Reynolds-Averaged Navier–Stokes (URANS) equations—used to numerically simulate approximate solutions for such turbulent flows. Rather than attempting to model each term in URANS separately, it is attractive to explore the capability of machine learning (ML) to model groups of terms, i.e. to directly model the force balances. We develop deep time-series ML for closure modeling of the URANS equations applied to SST. We consider decaying SST which are homogeneous and stably stratified by a uniform density gradient, enabling dimensionality reduction. We consider two time-series ML models: long short-term memory and neural ordinary differential equation. Both models perform accurately and are numerically stable in a posteriori (online) tests. Furthermore, we explore the data requirements of the time-series ML models by extracting physically relevant timescales of the complex system. We find that the ratio of the timescales of the minimum information required by the ML models to accurately capture the dynamics of the SST corresponds to the Reynolds number of the flow. The current framework provides the backbone to explore the capability of such models to capture the dynamics of high-dimensional complex dynamical system like SST flows.

97 MATHEMATICS AND COMPUTING

Understanding latent timescales in neural ordinary differential equation models of advection-dominated dynamical systems

The neural ordinary differential equation (ODE) framework has shown considerable promise in recent years in developing highly accelerated surrogate models for complex physical systems characterized by partial differential equations (PDEs). For PDE-based systems, state-of-the-art neural ODE strategies leverage a two-step procedure to achieve this acceleration: a nonlinear dimensionality reduction step provided by an autoencoder, and a time integration step provided by a neural-network based model for the resultant latent space dynamics (the neural ODE). This work explores the applicability of such autoencoder-based neural ODE strategies for PDEs in which advection terms play a critical role. More specifically, alongside predictive demonstrations, physical insight into the sources of model acceleration (i.e., how the neural ODE achieves its acceleration) is the scope of the current study. Such investigations are performed by quantifying the effects of both autoencoder and neural ODE components on latent system time-scales using eigenvalue analysis of dynamical system Jacobians. To this end, the sensitivity of various critical training parameters – de-coupled versus end-to-end training, latent space dimensionality, and the role of training trajectory length, for example – to both model accuracy and the discovered latent system timescales is quantified. Furthermore, this work specifically uncovers the key role played by the training trajectory length (the number of rollout steps in the loss function during training) on the latent system timescales: larger trajectory lengths correlate with an increase in limiting neural ODE time-scales, and optimal neural ODEs are found to recover the largest time-scales of the full-order (ground-truth) system. Demonstrations are performed across fundamentally different unsteady fluid dynamics configurations influenced by advection: (1) the Kuramoto–Sivashinsky equations (2) Hydrogen-Air channel detonations (the compressible reacting Navier–Stokes equations with detailed chemistry), and (3) 2D Atmospheric flow.

Advection-dominated dynamical systems

150 Shades of Green: Using the Full Spectrum of Remote Sensing Reflectance to Elucidate Color Shifts in the Ocean

This article proposes a simple and intuitive classification system by which to define full spectral remote sensing reflectance (Rrs(λ)) data with a quantitative output that enables a more manageable handling of spectral information for aquatic science applications. The weighted harmonic mean of the Rrs(λ) wavelengths outputs an Apparent Visible Wavelength (in units of nanometers), representing a one-dimensional geophysical metric of color that is inherently correlated to spectral shape. This dimensionality reduction of spectral information combined with the output along a continuum of wavelength values offers a robust and user-friendly means to describe and analyze spectral Rrs(λ) in terms of spatial and temporal trends and variability. The uncertainty in the algorithm's estimation of spectral shape is demonstrated on a global scale, in addition to the utility of the algorithm to discern spectral-spatial-temporal trends in the ocean, on a per-pixel basis for the entire 22 year continuous ocean color (SeaWiFS and MODIS-Aqua) time-series. This technique can be applied to datasets of varying multi- and hyper-spectral resolutions, providing continuity between heritage and future satellite sensors, and further enabling an effective means of elucidating similarities or differences in complex spectral signatures within the constraints of two dimensions. This straightforward means of conceptualizing multi-dimensional variability can help maximize the potential of the spectral information embedded in remote sensing data.

ocean color

Generative learning for slow manifolds and bifurcation diagrams

In dynamical systems characterized by separation of time scales, the approximation of so called “slow manifolds”, on which the long term dynamics lie, is a useful step for model reduction. Initializing on such slow manifolds is a useful step in modeling, since it circumvents fast transients, and is crucial in multiscale algorithms (like the equation-free approach) alternating between fine scale (fast) and coarser scale (slow) simulations. In a similar spirit, when one studies the infinite time dynamics of systems depending on parameters, the system attractors (e.g., its steady states) lie on bifurcation diagrams (curves for one-parameter continuation, and more generally, on manifolds in state parameter space. Sampling these manifolds gives us representative attractors (here, steady states of ODEs or PDEs) at different parameter values. Algorithms for the systematic construction of these manifolds (slow manifolds, bifurcation diagrams) are required parts of the “traditional” numerical nonlinear dynamics toolkit. In more recent years, as the field of Machine Learning develops, conditional score-based generative models (cSGMs) have been demonstrated to exhibit remarkable capabilities in generating plausible data from target distributions that are conditioned on some given label. It is tempting to exploit such generative models to produce samples of data distributions (points on a slow manifold, steady states on a bifurcation surface) conditioned on (consistent with) some quantity of interest (QoI, observable). In this work, we present a framework for using cSGMs to quickly (a) initialize on a low-dimensional (reduced-order) slow manifold of a multi-time-scale system consistent with desired value(s) of a QoI (a “label”) on the manifold, and (b) approximate steady states in a bifurcation diagram consistent with a (new, out-of-sample) parameter value. This conditional sampling can help uncover the geometry of the reduced slow-manifold and/or approximately “fill in” missing segments of steady states in a bifurcation diagram. Finally, the quantity of interest, which determines how the sampling is conditioned, is either known a priori or identified using manifold learning-based dimensionality reduction techniques applied to the training data.

Dynamical systems

Celestial Topology, Symmetry Theories, and Evidence for a NonSUSY D3‐Brane CFT

Symmetry Theories (SymThs) provide a flexible framework for analyzing the global categorical symmetries of a D -dimensional QFT D in terms of a (D + 1)-dimensional bulk system SymTh D+1 . In QFTs realized via local string backgrounds, these SymThs naturally arise from dimensional reduction of the linking boundary geometry. To track possible time dependent effects we introduce a celestial generalization of the standard “boundary at infinity” of a SymTh. As an application of these considerations we revisit large N quiver gauge theories realized by spacetime filling D3-branes probing a non-supersymmetric orbifold $\mathbb{R}$ 6 /Γ. Comparing the imprint of symmetry breaking on the celestial geometry at small and large ‘t Hooft coupling we find evidence for an intermediate symmetry preserving conformal fixed point.

conformal field theory

Bayesian Calibration of Stochastic Agent Based Model via Random Forest

Agent-based models (ABM) provide an excellent framework for modeling outbreaks and interventions in epidemiology by explicitly accounting for diverse individual interactions and environments. However, these models are usually stochastic and highly parametrized, requiring precise calibration for predictive performance. When considering realistic numbers of agents and properly accounting for stochasticity, this high-dimensional calibration can be computationally prohibitive. This paper presents a random forest-based surrogate modeling technique to accelerate the evaluation of ABMs and demonstrates its use to calibrate an epidemiological ABM named CityCOVID via Markov chain Monte Carlo (MCMC). The technique is first outlined in the context of CityCOVID's quantities of interest, namely hospitalizations and deaths, by exploring dimensionality reduction via temporal decomposition with principal component analysis (PCA) and via sensitivity analysis. The calibration problem is then presented, and samples are generated to best match COVID-19 hospitalization and death numbers in Chicago from March to June in 2020. Further, these results are compared with previous approximate Bayesian calibration (IMABC) results, and their predictive performance is analyzed, showing improved performance with a reduction in computation.

60 APPLIED LIFE SCIENCES

Toward memory-efficient melt pool monitoring: a classification framework using event-based imaging and sparse sensing technique

Vision sensors like CMOS and CCD cameras are often used for in-process monitoring of melt pools in laser-based additive and welding processes, but they require transferring large amounts of data and computational processing resources. Event-based neuromorphic imagery, on the other hand, detects only the change in pixel intensity, thus potentially reducing the data amount and latency. With an event imager, this study develops a framework for melt pool condition classification, including image construction, time scale selection, optimal pixel selection, and sparse classification, to achieve a highly memory-efficient scheme. These are based on sparse sensing techniques with singular value decomposition (SVD) and QR pivoting, the two fundamental matrix transformations for linear dimensionality reduction. The framework is then validated by classifying a controlled experiment by exciting various mode shapes of liquid gallium pools of varying depths (3, 6, and 8 mm). At 200 pixels, the classifier can reach overall accuracy of 75%, while at 2000 pixels (0.013% of the total possible pixels), the accuracy is nearly 90% (89.86%). At the same number of pixels, random selection can only achieve 46% and 67%, respectively. The memory savings of the sparsely sampled event data compared to a conventional imager is about 500 times. In addition to performance, implementation and limitations of the framework are also discussed.

42 ENGINEERING

Quantifying Microstructure Variability in Laser Powder Bed Fusion 316 L Stainless Steel Microstructures with Spatial Statistics

Here, we have explored data-driven methods for material microstructure quantification that improve sensitivity to microstructural changes compared to traditional approaches. The methods integrate multiple microstructural properties, including grain morphology, crystallographic orientation, and material phase information. The simpler method employs maps of the Euclidean distance transformation metric to evaluate the morphology of grain boundary networks. The more intensive approach employs generalized spherical harmonic mapping for crystallographic orientations, per-pixel phase information, and a variational auto-encoder for dimensionality reduction and results in a multidimensional clustering of by microstructure similarity. Applied to an experimental dataset of additively manufactured steel, both methods detected slight variations in samples produced under nominally identical processing conditions. Both methods were able to distinguish between samples from multiple (nominally identical) builds, while the generalized spherical harmonics-based method could additionally cluster data samples rotated at two orientations on the build plate. The improved sensitivity of the methods, demonstrated through comparison with traditional microstructure characterization techniques, offers advantages for microstructure quantification and comparisons in advanced manufacturing applications.

SS316L

Projection-based multifidelity linear regression for data-scarce applications

Surrogate modeling for systems with high-dimensional quantities of interest remains challenging, particularly when training data are costly to acquire. This work develops multifidelity methods for multiple-input multiple-output linear regression targeting data-limited applications with high-dimensional outputs. Multifidelity methods integrate many inexpensive low-fidelity model evaluations with limited, costly high-fidelity evaluations. We introduce two projection-based multifidelity linear regression approaches with linear and nonlinear features that leverage principal component basis vectors for dimensionality reduction and combine multifidelity data through: (i) a direct data augmentation using low-fidelity data, and (ii) a data augmentation incorporating explicit linear corrections between low-fidelity and high-fidelity data. The data augmentation approaches combine high-fidelity and low-fidelity data into a unified training set and train the linear regression model through weighted least squares with fidelity-specific weights. We introduce a proximity-based weighting scheme with automatic weight selection strategy through cross-validation. Here, the proposed multifidelity linear regression methods are demonstrated on approximating the surface pressure field of a hypersonic vehicle in flight and the temperature field on an aircraft disc braking system. In an ultra low-data regime of no more than twelve high-fidelity samples, multifidelity linear regression achieves approximately 2% – 12% improvement in median accuracy and a higher R 2 score relative to single-fidelity methods at comparable computational cost.

data augmentation

Karhunen–Loève deep learning method for surrogate modeling and approximate Bayesian parameter estimation

We evaluate the performance of the Karhunen-Loève Deep Neural Network (KL-DNN) framework for surrogate modeling and approximate Bayesian parameter estimation in partial differential equation models. In the surrogate model, the Karhunen-Loève (KL) expansions are used for the dimensionality reduction of the number of unknown parameters and variables, and a deep neural network is employed to relate the reduced space of parameters to that of the state variables. The KL-DNN surrogate model is used to formulate a maximum-a-posteriori-like least-squares problem, which is randomized to draw samples of the posterior distribution of the parameters. We test the proposed framework for a hypothetical unconfined aquifer via comparison with the forward MODFLOW and inverse PEST++ iterative ensemble smoother (IES) solutions as well as the state-of-the-art Fourier neural operator (FNO) and deep operator networks (DeepONets) operator learning surrogate models. Our results show that the KL-DNN surrogate model outperforms FNO and DeepONet for forward predictions. For solving inverse problems, the randomized algorithm provides the same or more accurate Bayesian predictions of the parameters than IES as evidenced by the higher log-predictive probability of both the estimated parameter field and the forecast hydraulic head. The posterior mean obtained from the randomized algorithm is closer to the reference parameter field than that obtained with FNO as the maximum a posteriori estimate.

Approximate Bayesian inference

Personalized and uncertainty-aware coronary hemodynamics simulations: From Bayesian estimation to improved multi-fidelity uncertainty quantification

Non-invasive simulations of coronary hemodynamics have improved clinical risk stratification and treatment outcomes for coronary artery disease, compared to relying on anatomical imaging alone. However, simulations typically use empirical approaches to distribute total coronary flow amongst the arteries in the coronary tree, which ignores patient variability, the presence of disease, and other clinical factors. Further, uncertainty in the clinical data often remains unaccounted for in the modeling pipeline. We present an end-to-end uncertainty-aware pipeline to (1) personalize coronary flow simulations by incorporating vessel-specific coronary flows as well as cardiac function; and (2) predict clinical and biomechanical quantities of interest with improved precision, while accounting for uncertainty in the clinical data. We assimilate patient-specific measurements of myocardial blood flow from clinical CT myocardial perfusion imaging to estimate branch-specific coronary artery flows. Simulated noise in the clinical data is used to estimate the joint posterior distributions of the model parameters using adaptive Markov Chain Monte Carlo sampling. Additionally, the posterior predictive distribution for the relevant quantities of interest is determined using a new approach combining multi-fidelity Monte Carlo estimation with non-linear, data-driven dimensionality reduction. This leads to improved correlations between high- and low-fidelity model outputs. Our framework accurately recapitulates clinically measured cardiac function as well as branch-specific coronary flows under measurement noise uncertainty. We observe substantial reductions in confidence intervals for estimated quantities of interest compared to single-fidelity Monte Carlo estimation and state-of-the-art multi-fidelity Monte Carlo methods. This holds especially true for quantities of interest that showed limited correlation between the low- and high-fidelity model predictions. In addition, the proposed multi-fidelity Monte Carlo estimators are significantly cheaper to compute than traditional estimators, under a specified confidence level or variance. The proposed pipeline for personalized and uncertainty-aware predictions of coronary hemodynamics is based on routine clinical measurements and recently developed techniques for CT myocardial perfusion imaging. The proposed pipeline offers significant improvements in precision and reduction in computational cost.

Bayesian parameter estimation

Online task-space motion control for positioner-coordinated multi-robot manufacturing systems

Incorporating multiple robotic manipulators into large-scale manufacturing systems enhances production efficiency and expands manufacturing capabilities beyond those of single-robot systems. Workpiece positioners in robotic manufacturing have demonstrated significant benefits for process optimization, but coordination strategies for multi-robot systems with shared positioners have received limited attention. This work presents a task-space coordinated trajectory-tracking control framework for multi-robot manufacturing systems, in which robots coordinate their motions within a shared, dynamic workpiece positioning frame. A workpiece positioner actively adjusts the pose of the manufactured component to enable greater operational concurrency and improve overall production efficiency. The proposed motion-coordination scheme employs a distributed and scalable architecture, supporting coordination across heterogeneous multi-robot systems. Two optimization methodologies are introduced to manage kinematic redundancies and maintain continuous, near-optimal operation throughout the manufacturing process. The first strategy exploits a task-space dimensionality reduction to achieve locally optimal configurations by leveraging symmetry-axis rotations of the tool. The second strategy utilizes the workpiece positioner to drive the coordinated robots toward stable and kinematically favorable configurations. For both optimization strategies, multiple objectives are defined to improve key performance metrics, including manipulability, configuration consistency, proximity to mechanical limits, and motion efficiency. Addressing a key limitation of existing coordination approaches, the framework is designed around online setpoint modification, allowing coordinated robots to respond effectively to in-situ process feedback. The proposed control framework is validated using the Robot Operating System (ROS) middleware on a combination of physical and simulated multi-robot system hardware.

Arbogast, Alex [ORNL] (ORCID:0000000154740723)

Navigating Large Chemical Spaces Using Graph Theory and Integer Programming

Navigating and analyzing large chemical spaces are necessary to accelerate the design and discovery of new molecules and chemical processes. In this work, we introduce a computational framework that integrates graph theory and integer programming to enable the efficient navigation of large chemical spaces. Our framework represents the chemical space as a graph, wherein nodes represent molecules and edges represent the degree of similarity or connectivity based on domain-specific information. Using the graph representation, we identify representative molecules by computing the so-called minimum dominating set (MDS), which in our context is the minimum set of molecules that is connected to all other molecules. We present a suite of solution strategies for the MDS problem including heuristic and rigorous integer programming (IP) approaches. We show that these approaches allow us to capture physicochemical properties and domain-specific logic and constraints, facilitating the identification of molecules with the target properties. We demonstrate the effectiveness of the proposed approach by navigating the chemical space of per- and polyfluoroalkyl substances (PFAS); this comprises approximately 15,000 molecular structures. We compare our framework against traditional dimensionality reduction and clustering methods such as t-SNE and K-means clustering.

Chemical structure

Reduced‐Order Probabilistic Emulation of Physics‐Based Ring Current Models: Application to RAM‐SCB Particle Flux

Abstract In this work, we address the computational challenge of large‐scale physics‐based simulation models for the ring current. Reduced computational cost allows for significantly faster than real‐time forecasting, enhancing our ability to predict and respond to dynamic changes in the ring current, valuable for space weather monitoring and mitigation efforts. Additionally, it can also be used for a comprehensive investigation of the system. Thus, we aim to create an emulator for the Ring current‐Atmosphere interactions Model with Self‐Consistent magnetic field (RAM‐SCB) particle flux that not only improves efficiency but also facilitates forecasting with reliable estimates of prediction uncertainties. The probabilistic emulator is built upon the methodology developed by Licata and Mehta (2023), https://doi.org/10.1029/2022sw003345 . A novel discrete sampling is used to identify 30 simulation periods over 20 years of solar and geomagnetic activity. Focusing on a subset of particle flux, we use Principal Component Analysis for dimensionality reduction and Long Short‐Term Memory (LSTM) neural networks to perform dynamic modeling. Hyperparameter space was explored extensively resulting in about 5% median symmetric accuracy across all data sets for one‐step dynamic prediction. Using a hierarchical ensemble of LSTMs, we have developed a reduced‐order probabilistic emulator (ROPE) tailored for time‐series forecasting of particle flux in the ring current. This ROPE offers accurate predictions of omnidirectional flux at a single energy with no pitch angle information, providing robust predictions on the test set with an error score below 11% and calibration scores under 8% with bias under 2% providing a significant speed up as compared to the full RAM‐SCB run.

79 ASTRONOMY AND ASTROPHYSICS

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry