Search NASA⌕ Search

SEARCH · Search NASA

Results for “Many body expansion”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Delocalization error poisons the density-functional many-body expansion

The many-body expansion is a fragment-based approach to large-scale quantum chemistry that partitions a single monolithic calculation into manageable subsystems. This technique is increasingly being used as a basis for fitting classical force fields to electronic structure data, especially for water and aqueous ions, and for machine learning. Here, we show that the many-body expansion based on semilocal density functional theory affords wild oscillations and runaway error accumulation for ion–water interactions, typified by F − (H 2 O) N with N ≳ 15. We attribute these oscillations to self-interaction error in the density-functional approximation. The effect is minor or negligible in small water clusters, explaining why it has not been noticed previously, but grows to catastrophic proportion in clusters that are only moderately larger. This behavior can be counteracted with hybrid functionals but only if the fraction of exact exchange is ≳50%, whereas modern meta-generalized gradient approximations including ωB97X-V, SCAN, and SCAN0 are insufficient to eliminate divergent behavior. Other mitigation strategies including counterpoise correction, density correction (i.e., exchange–correlation functionals evaluated atop Hartree–Fock densities), and dielectric continuum boundary conditions do little to curtail the problematic oscillations. In contrast, energy-based screening to cull unimportant subsystems can successfully forestall divergent behavior. These results suggest that extreme caution is warranted when the many-body expansion is combined with density functional theory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Untangling Sources of Error in the Density-Functional Many-Body Expansion

The many-body expansion provides a framework for data-driven applications of electronic structure theory, including parametrization of classical force fields and machine learning. In this article, we demonstrate that its use significantly amplifies quadrature grid errors when modern density-functional approximations are employed. Standard grids that work well in conventional density-functional calculations result in runaway error accumulation when used with the many-body expansion. At the same time, delocalization error is also exacerbated, leading to exaggerated estimates of nonadditive n-body interactions. This is illustrated for anion–water clusters using the SCAN, r2SCAN, ωB97X-V and ωB97M-V functionals. By employing dense quadrature grids, the inherent self-interaction error is exposed, which can then be mitigated using a variety of other strategies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Hierarchical Truncations for Many-Body Expansion Potentials

In this work, a new strategy to truncate high-order terms in the many-body expansion (MBE) is proposed. This new approach, which we call a hierarchical many-body expansion (HMBE), is based on a hierarchical partition of the system into multitier fragments and can in principle be applied to any large molecular system. Numerical tests on a series of (H 2 O) 64 structures are presented, demonstrating satisfactory relative energies between the clusters and binding energies of individual clusters compared with full-cluster calculations, with significantly fewer high-order terms computed than conventional MBE. The hierarchical truncation can be augmented by certain many-body terms for fragments at the interface between the partitions (called “Schengen terms”) to further improve accuracy. This work establishes the HMBE scheme as a promising framework to model very large systems (e.g., proteins), which are naturally built on a hierarchical structure.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable generalized screening for high-order terms in the many-body expansion: Algorithm, open-source implementation, and demonstration

The many-body expansion lies at the heart of numerous fragment-based methods that are intended to sidestep the nonlinear scaling of ab initio quantum chemistry, making electronic structure calculations feasible in large systems. In principle, inclusion of higher-order n-body terms ought to improve the accuracy in a controllable way, but unfavorable combinatorics often defeats this in practice and applications with n ≥ 4 are rare. Here, we outline an algorithm to overcome this combinatorial bottleneck, based on a bottom-up approach to energy-based screening. This is implemented within a new open-source software application (“Fragme∩t”), which is integrated with a lightweight semi-empirical method that is used to cull subsystems, attenuating the combinatorial growth of higher-order terms in the graph that is used to manage the calculations. This facilitates applications of unprecedented size, and we report four-body calculations in (H2O)64 clusters that afford relative energies within 0.1 kcal/mol/monomer of the supersystem result using less than 10% of the unique subsystems. We also report n-body calculations in (H2O)20 clusters up to n = 8, at which point the expansion terminates naturally due to screening. These are the largest n-body calculations reported to date using ab initio electronic structure theory, and they confirm that high-order n-body terms are mostly artifacts of basis-set superposition error.

Chemistry↗

Energy-Screened Many-Body Expansion for Protein–Ligand Interactions: Examining Convergence for Metalloenzymes Through Seven–Body Interactions

Fragment-based quantum chemistry is a powerful strategy for calculating protein−ligand interaction energies using quantum chemistry methods. Rigorous convergence often requires hundreds of atoms in the protein binding-site model, especially if that model is constructed using distance-based criteria to select amino acid residues, while three- and four-body calculations exhibit instability related to combinatorial proliferation in the number of subsystem calculations. Here, we report an energy-based screening protocol for the many-body expansion applied to protein−ligand interactions, implemented in the open-source FRAGME∩T code. Using a combination of aggressive screening based on semiempirical quantum chemistry, with an improved graph-theoretical algorithm to eliminate unimportant subsystems, we are able to perform n-body calculations up to n = 7 using density functional theory in triple-ζ basis sets. Distance cutoffs further reduce the cost without compromising accuracy. Rapid and stable convergence of the many-body expansion is obtained by n = 4, for a pair of metalloenzymes in which a divalent ion coordinates directly to the ligand. As compared to previous results that relied solely on distance cutoffs, oscillations in the n-body corrections are reduced or eliminated, although residual errors remain in one case. This work demonstrates that benchmark-quality protein−ligand interaction energies can be systematically converged using a method with excellent parallel efficiency and scalability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Many-Body Expansion for Metals: II. Nonadditive Terms in Clusters Composed of Metals with n s 1 , n s 2 , and n s 2 p 1 Configurations

The many-body expansion (MBE) was applied to homometallic and heterometallic trimers of metals with ns 1 , ns 2 , and ns 2 p 1 configurations to investigate its convergence, the magnitude and nature (stabilizing/destabilizing) of the individual terms and seek an understanding of their variation across the different families of clusters. In particular, we examined the series of alkali metals (Li 3 , Na 3 , K 3 , Li 2 Na, LiNa 2 ), alkali metal borides (Li 2 B and LiB 2 ), and alkaline earth metals (Be 3 , Be 2 Mg, BeMg 3 , and Mg 3 ) trimers, as well as sodium clusters Na n , n = 2–5. Here, we found that there is no uniform contribution (stabilizing or destabilizing) across the series in the different families of trimers. For instance, the 2-B term stabilizes the ground states of the Na 3 (doublet), Na 4 (singlet), and Na 5 (doublet) clusters, and the 3-B term destabilizes them; however, the opposite holds for the quartet state of the Li 3 , Li 2 Na, LiNa 2 , and Na 3 clusters (destabilizing 2-B, stabilizing 3-B). Substituting Li with B in the quartet state of Li 3 results in a significant reduction of the 3-B term amounting to 16% (Li 2 B) and 5% (LiB 3 ) of the binding energy. On the contrary, the ground states of the alkaline earth metal clusters (Be 3 , Be 2 Mg, BeMg 3 , and Mg 3 ) are stabilized by the 3-B term, while the 2-B term destabilizes them. Overall, we find that the 3-B terms significantly stabilize the high-spin multiplicity states of the ns 1 configurations and the low-spin states of the ns 2 configurations. Finally, as the size of the metal increases, the contribution of the 3-B term to the binding energy decreases due to the longer metal–metal bond distances.

Alkali metals↗

Ligand Many-Body Expansion as a General Approach for Accelerating Transition Metal Complex Discovery

Methods that accelerate the evaluation of molecular properties are essential for chemical discovery. While some degree of ligand additivity has been established for transition metal complexes, it is underutilized in asymmetric complexes, such as the square pyramidal coordination geometries highly relevant to catalysis. To develop predictive methods beyond simple additivity, we apply a many-body expansion to octahedral and square pyramidal complexes and introduce a correction based on adjacent ligands (i.e., the cis interaction model). We first test the cis interaction model on adiabatic spin-splitting energies of octahedral Fe(II) complexes, predicting DFT-calculated values of unseen binary complexes to within an average of 1.4 kcal/mol. Uncertainty analysis reveals the optimal basis, comprising the homoleptic and mer symmetric complexes. We next show that the cis model (i.e., the cis interaction model solved for the optimal basis) infers both DFT- and CCSD(T)-calculated model catalytic reaction energies to within 1 kcal/mol on average. The cis model predicts low-symmetry complexes with reaction energies outside the range of binary complex reaction energies. We observe that trans interactions are unnecessary for most monodentate systems but can be important for some combinations of ligands, such as complexes containing a mixture of bidentate and monodentate ligands. Lastly, we demonstrate that the cis model may be combined with Δ-learning to predict CCSD(T) reaction energies from exhaustively calculated DFT reaction energies and the same fraction of CCSD(T) reaction energies needed for the cis model, achieving around 30% of the error from using the CCSD(T) reaction energies in the cis model alone.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)↗

Erratum: “Breaking covalent bonds in the context of the many-body expansion (MBE). I. The purported ‘first row anomaly’ in XH n (X = C, Si, Ge, Sn; n = 1–4)” [J. Chem. Phys. 156, 244303 (2022)]

We have noted typographical errors in TABLE V of J. Chem. Phys. 156, 244303 (2022). Specifically, the values of the angles φ HχH for the XH 2 species in both the ( 3 B 1 ) and ( 1 A 1 ) states were incorrectly reported as half of the correct values. Additionally, the values of the same angles for the XH 3 and XH 4 species were reported correctly.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Many-Body Perspective of Nuclear Quantum Effects in Aqueous Clusters

Nuclear quantum effects play an important role in the structure and thermodynamics of aqueous systems. Here, by performing a many-body expansion with nuclear-electronic orbital (NEO) theory, we show that proton quantization can give rise to significant energetic contributions for many-body interactions spanning several molecules in single-point energy calculations of water clusters. Although zero-point motion produces a large increase in energy at the one-body level, nuclear quantum effects serve to stabilize higher-order molecular interactions. These results are significant because they demonstrate that nuclear quantum effects play a nontrivial role in many-body interactions of aqueous systems. Our approach also provides a pathway for incorporating nuclear quantum effects into water potential energy surfaces. The NEO approach is advantageous for many-body expansion analyses because it includes nuclear quantum effects directly in the energies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Many-Body Basis Set Amelioration Method for Incremental Full Configuration Interaction

Incremental full configuration interaction (iFCI) is a polynomial-cost electronic structure method that systematically approaches the FCI limit by employing the method of increments to solve the Schrödinger equation through a many-body expansion. This article introduces the many-body basis set amelioration (MBBSA) method, which is designed to allow iFCI to be applicable to larger atomic orbital basis sets. MBBSA uses a series of inexpensive iFCI calculations to approximate the correlation energy that would be found using a more expensive, highly accurate iFCI calculation. Here, when compared to standard iFCI computations on smaller molecules in triple-zeta and larger basis sets, MBBSA provides approximations to the total and relative energies within chemical accuracy. MBBSA exhibits a reduced cost of between 60-92% when compared to standard iFCI calculations, with larger systems experiencing the largest benefit. Tests of MBBSA on two reactions that involve highly correlated systems, the automerization of cyclobutadiene and a Criegee intermediate reaction, show that MBBSA has practical utility for studying realistic chemistries.

Basis sets↗

Force Field X: A computational microscope to study genetic variation and organic crystals using theory and experiment

Force Field X (FFX) is an open-source software package for atomic resolution modeling of genetic variants and organic crystals that leverages advanced potential energy functions and experimental data. FFX currently consists of nine modular packages with novel algorithms that include global optimization via a many-body expansion, acid–base chemistry using polarizable constant-pH molecular dynamics, estimation of free energy differences, generalized Kirkwood implicit solvent models, and many more. Applications of FFX focus on the use and development of a crystal structure prediction pipeline, biomolecular structure refinement against experimental datasets, and estimation of the thermodynamic effects of genetic variants on both proteins and nucleic acids. The use of Parallel Java and OpenMM combines to offer shared memory, message passing, and graphics processing unit parallelization for high performance simulations. Overall, the FFX platform serves as a computational microscope to study systems ranging from organic crystals to solvated biomolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Pitfalls in the n -mode representation of vibrational potentials

Simulations of anharmonic vibrational motion rely on computationally expedient representations of the governing potential energy surface. The n-mode representation (n-MR)—effectively a many-body expansion in the space of molecular vibrations—is a general and efficient approach that is often used for this purpose in vibrational self-consistent field (VSCF) calculations and correlated analogues thereof. In the present analysis, a lack of convergence in many VSCF calculations is shown to originate from negative and unbound potentials at truncated orders of the n-MR expansion. For cases of strong anharmonic coupling between modes, the n-MR can both dip below the true global minimum of the potential surface and lead to effective single-mode potentials in VSCF that do not correspond to bound vibrational problems, even for bound total potentials. The present analysis serves mainly as a pathology report of this issue. Furthermore, this insight into the origin of VSCF non-convergence provides a simple, albeit ad hoc, route to correct the problem by “painting in” the full representation of groups of modes that exhibit these negative potentials at little additional computational cost. Somewhat surprisingly, this approach also reasonably approximates the results of the next-higher n-MR order and identifies groups of modes with particularly strong coupling. The method is shown to identify and correct problematic triples of modes—and restore SCF convergence—in two-mode representations of challenging test systems, including the water dimer and trimer, as well as protonated tropine.

Chemistry↗

Taming the virtual space for incremental full configuration interaction

Incremental full configuration interaction (iFCI) closely approximates the FCI limit with polynomial cost through a many-body expansion of the correlation energy, providing highly accurate total energies within a given basis set. To extend iFCI beyond previous basis set limitations, this work introduces a novel natural orbital (NO) screening approach, incremental NO full configuration interaction (iNO-FCI). By consideration of the importance of virtual orbital selection in the convergence of iFCI, iNO-FCI maximizes the consistency between orbitals selected for each correlated body. iNO-FCI employs a principle of cancellation of errors and ensures that the same set of virtual NOs is used for interdependent terms. Here, this strategy significantly reduces computational cost without compromising precision. Computational savings of up to 95% are demonstrated, allowing access to larger basis sets that were previously computationally prohibitive. iNO-FCI is herein introduced and benchmarked for several difficult test cases involving double-bond dissociation, biradical systems, conjugated π systems, and the spin gap of a Cu-based transition metal complex.

Correlation energy↗

Quick-and-Easy Validation of Protein–Ligand Binding Models Using Fragment-Based Semiempirical Quantum Chemistry

Electronic structure calculations in enzymes converge very slowly with respect to the size of the model region that is described using quantum mechanics (QM), requiring hundreds of atoms to obtain converged results and exhibiting substantial sensitivity (at least in smaller models) to which amino acids are included in the QM region. As such, there is considerable interest in developing automated procedures to construct a QM model region based on well-defined criteria. However, testing such procedures is burdensome due to the cost of large-scale electronic structure calculations. Here, we show that semiempirical methods can be used as alternatives to density functional theory (DFT) to assess convergence in sequences of models generated by various automated protocols. The cost of these convergence tests is reduced even further by means of a many-body expansion. We use this approach to examine convergence (with respect to model size) of protein–ligand binding energies. Fragment-based semiempirical calculations afford well-converged interaction energies in a tiny fraction of the cost required for DFT calculations. Two-body interactions between the ligand and single-residue amino acid fragments afford a low-cost way to construct a “QM-informed” enzyme model of reduced size, furnishing an automatable active-site model-building procedure. This provides a streamlined, user-friendly approach for constructing ligand binding-site models that needs neither a priori information nor manual adjustments. Extension to model-building for thermochemical calculations should be straightforward.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Convergent Protocols for Computing Protein–Ligand Interaction Energies Using Fragment-Based Quantum Chemistry

Fragment-based quantum chemistry methods offer a way to sidestep the steep nonlinear scaling of electronic structure calculations so that large molecular systems can be investigated using high-level methods. Here, we use fragmentation to compute protein–ligand interaction energies in systems with several thousand atoms, using a new software platform for managing fragment-based calculations that implements a screened many-body expansion. Convergence tests using a minimal-basis semiempirical method (HF-3c) indicate that two-body calculations, with single-residue fragments and simple hydrogen caps, are sufficient to reproduce interaction energies obtained using conventional supramolecular electronic structure calculations, to within 1 kcal/mol at about 1% of the computational cost. We also demonstrate that the HF-3c results are illustrative of trends obtained with density functional theory in basis sets up to augmented quadruple-ζ quality. Strategic deployment of fragmentation facilitates the use of converged biomolecular model systems alongside high-quality electronic structure methods and basis sets, bringing ab initio quantum chemistry to systems of hitherto unimaginable size. This will be useful for generation of high-quality training data for machine learning applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fragme∩t: An Open‐Source Framework for Multiscale Quantum Chemistry Based on Fragmentation

Fragment-based quantum chemistry offers a means to circumvent the nonlinear computational scaling of conventional electronic structure calculations, by partitioning a large calculation into smaller subsystems then considering the many-body interactions between them. Variants of this approach have been used to parameterize classical force fields and machine learning potentials, applications that benefit from interoperability between quantum chemistry codes. However, there is a dearth of software that provides interoperability yet is purpose-built to handle the combinatorial complexity of fragment-based calculations. To fill this void we introduce “Fragme∩t”, an open-source software application that provides a tool for community validation of fragment-based methods, a platform for developing new approximations, and a framework for analyzing many-body interactions. Fragme∩t includes algorithms for automatic fragment generation and structure modification, and for distance- and energy-based screening of the requisite subsystems. Checkpointing, database management, and parallelization are handled internally and results are archived in a portable database. Interfaces to various quantum chemistry engines are easy to write and exist already for Q-Chem, PySCF, xTB, Orca, CP2K, MRCC, Psi4, NWChem, GAMESS, and MOPAC. Applications reported here demonstrate parallel efficiencies around 96% on more than 1000 processors but also showcase that the code can handle large-scale protein fragmentation using only workstation hardware, all with a codebase that is designed to be usable by non-experts. Fragme∩t conforms to modern software engineering best practices and is built upon well established technologies including Python, SQLite, and Ray. The source code is available under the Apache 2.0 license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗