Search NASA⌕ Search

Engineering topics

Datta, Dipayan

Publications and source records attributed to Datta, Dipayan.

The General Atomic and Molecular Electronic Structure System (GAMESS): Novel Methods on Novel Architectures

The primary focus of GAMESS over the last 5 years has been the development of new high-performance codes that are able to take effective and efficient advantage of the most advanced computer architectures, both CPU and accelerators. These efforts include employing density fitting and fragmentation methods to reduce the high scaling of well-correlated (e.g., coupled-cluster) methods as well as developing novel codes that can take optimal advantage of graphical processing units and other modern accelerators. Because accurate wave functions can be very complex, an important new functionality in GAMESS is the quasi-atomic orbital analysis, an unbiased approach to the understanding of covalent bonds embedded in the wave function. Finally, best practices for the maintenance and distribution of GAMESS are also discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Porting fragmentation methods to GPUs using an OpenMP API: Offloading the resolution-of-the-identity second-order Møller–Plesset perturbation method

Here, using an OpenMP Application Programming Interface, the resolution-of-the-identity second-order Møller–Plesset perturbation (RI-MP2) method has been off-loaded onto graphical processing units (GPUs), both as a standalone method in the GAMESS electronic structure program and as an electron correlation energy component in the effective fragment molecular orbital (EFMO) framework. First, a new scheme has been proposed to maximize data digestion on GPUs that subsequently linearizes data transfer from central processing units (CPUs) to GPUs. Second, the GAMESS Fortran code has been interfaced with GPU numerical libraries (e.g., NVIDIA cuBLAS and cuSOLVER) for efficient matrix operations (e.g., matrix multiplication, matrix decomposition, and matrix inversion). The standalone GPU RI-MP2 code shows an increasing speedup of up to 7.5× using one NVIDIA V100 GPU with one IBM 42-core P9 CPU for calculations on fullerenes of increasing size from 40 to 260 carbon atoms using the 6-31G(d)/cc-pVDZ-RI basis sets. A single Summit node with six V100s can compute the RI-MP2 correlation energy of a cluster of 175 water molecules using the correlation consistent basis sets cc-pVDZ/cc-pVDZ-RI containing 4375 atomic orbitals and 14 700 auxiliary basis functions in ~0.85 h. In the EFMO framework, the GPU RI-MP2 component shows near linear scaling for a large number of V100s when computing the energy of an 1800-atom mesoporous silica nanoparticle in a bath of 4000 water molecules. The parallel efficiencies of the GPU RI-MP2 component with 2304 and 4608 V100s are 98.0% and 96.1%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PDG: A Composite Method Based on the Resolution of the Identity

The Gaussian-3 (G3) composite approach for thermochemical properties is revisited in light of the enhanced computational efficiency and reduced memory costs by applying the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs) to the computationally demanding component methods in the G3 model: the energy and gradient computations via the second-order Møller–Plesset perturbation theory (MP2) and the energy computations using the coupled-cluster singles–doubles method augmented with noniterative triples corrections [CCSD(T)]. Efficient implementation of the RI-based methods is achieved by employing a hybrid distributed/shared memory model based on MPI and OpenMP. The new variant of the G3 composite approach based on the RI approximation is termed the RI-G3 scheme, or alternatively the PDG method. The accuracy of the new RI-G3/PDG scheme is compared to the “standard” G3 composite approach that employs the memory-expensive four-center ERIs in the MP2 and CCSD(T) calculations. Taking the computation of the heats of formation of the closed-shell molecules in the G3/99 test set as a test case, it is demonstrated that the RI approximation introduces negligible changes to the mean absolute errors relative to the standard G3 model (less than 0.1 kcal/mol), while the standard deviations remain unaltered. Furthermore, the efficiency and memory requirements for the RI-MP2 and RI-CCSD(T) methods are compared to the standard MP2 and CCSD(T) approaches, respectively. The hybrid MPI/OpenMP-based RI-MP2 energy plus gradient computation is found to attain a 7.5× speedup over the standard MP2 calculations. For the most demanding CCSD(T) calculations, the application of the RI approximation is found to nearly halve the memory demand, confer about a 4–5× speedup for the CCSD iterations, and reduce the computational time for the compute-intensive triples correction step by several hours.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Recent Developments in the General Atomic and Molecular Electronic Structure System

A discussion of many of the recently implemented features of GAMESS (General Atomic and Molecular Electronic Structure System) and LibCChem (the C++ CPU/GPU library associated with GAMESS) is presented. These features include fragmentation methods like the fragment molecular orbital, effective fragment potential and effective fragment molecular orbital methods, hybrid MPI/OpenMP approaches to Hartree-Fock and resolution of the identity second order perturbation theory. Many new coupled cluster theory methods have been implemented in GAMESS, as have multiple levels of density functional/tight binding theory. The role of accelerators, especially graphical processing units, is discussed in the context of the new features of LibCChem, as is the associated problem of power consumption as the power of computers increases dramatically. The process by which a complex program suite like GAMESS is maintained and developed is considered. Future developments are briefly summarized.

74 ATOMIC AND MOLECULAR PHYSICS↗