Search NASA⌕ Search

Engineering topics

Yu, Victor Wen-zhe

Publications and source records attributed to Yu, Victor Wen-zhe.

Molecular NMR shieldings, J -couplings, and magnetizabilities from numeric atom-centered orbital based density-functional calculations

This paper reports and benchmarks a new implementation of nuclear magnetic resonance shieldings, magnetizabilities, and J-couplings for molecules within semilocal density functional theory, based on numeric atom-centered orbital (NAO) basis sets. NAO basis sets are attractive for the calculation of these nuclear magnetic resonance (NMR) parameters because NAOs provide accurate atomic orbital representations especially near the nucleus, enabling high-quality results at modest computational cost. Moreover, NAOs are readily adaptable for linear scaling methods, enabling efficient calculations of large systems. Here, the paper has five main parts: (1) It reviews the formalism of density functional calculations of NMR parameters in one comprehensive text to make the mathematical background available in a self-contained way. (2) The paper quantifies the attainable precision of NAO basis sets for shieldings in comparison to specialized Gaussian basis sets, showing similar performance for similar basis set size. (3) The paper quantifies the precision of calculated magnetizabilities, where the NAO basis sets appear to outperform several established Gaussian basis sets of similar size. (4) The paper quantifies the precision of computed J-couplings, for which a group of customized NAO basis sets achieves precision of ~Hz for smaller basis set sizes than some established Gaussian basis sets. (5) The paper demonstrates that the implementation is applicable to systems beyond 1000 atoms in size.

74 ATOMIC AND MOLECULAR PHYSICS↗

First-Principles Investigation of Near-Surface Divacancies in Silicon Carbide

The realization of quantum sensors using spin defects in semiconductors requires a thorough understanding of the physical properties of the defects in the proximity of surfaces. We report a study of the divacancy (V Si V C ) in 3C-SiC, a promising material for quantum applications, as a function of surface reconstruction and termination with -H, -OH, -F and oxygen groups. Here we show that a V Si V C close to hydrogen-terminated (2 x 1) surfaces is a robust spin-defect with a triplet ground state and no surface states in the band gap and with small variations of many of its physical properties relative to the bulk, including the zero-phonon line and zero-field splitting. However, the Debye-Waller factor decreases in the vicinity of the surface and our calculations indicate it may be improved by strain-engineering. Overall our results show that the V Si V C close to SiC surfaces is a promising spin defect for quantum applications, similar to its bulk counterpart.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Excited State Properties of Point Defects in Semiconductors and Insulators Investigated with Time-Dependent Density Functional Theory

Here, we present a formulation of spin-conserving and spin-flip hybrid time-dependent density functional theory (TDDFT), including the calculation of analytical forces, which allows for efficient calculations of excited state properties of solid-state systems with hundreds to thousands of atoms. We discuss an implementation on both GPU- and CPU-based architectures along with several acceleration techniques. We then apply our formulation to the study of several point defects in semiconductors and insulators, specifically the negatively charged nitrogen-vacancy and neutral silicon-vacancy centers in diamond, the neutral divacancy center in 4H silicon carbide, and the neutral oxygen-vacancy center in magnesium oxide. Our results highlight the importance of taking into account structural relaxations in excited states in order to interpret and predict optical absorption and emission mechanisms in spin defects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GPU Acceleration of Large-Scale Full-Frequency GW Calculations

Many-body perturbation theory is a powerful method to simulate electronic excitations in molecules and materials starting from the output of density functional theory calculations. By implementing the theory efficiently so as to run at scale on the latest leadership high-performance computing systems it is possible to extend the scope of GW calculations. Here, we present a GPU acceleration study of the full-frequency GW method as implemented in the WEST code. Excellent performance is achieved through the use of (i) optimized GPU libraries, e.g., cuFFT and cuBLAS, (ii) a hierarchical parallelization strategy that minimizes CPU-CPU, CPU-GPU, and GPU-GPU data transfer operations, (iii) nonblocking MPI communications that overlap with GPU computations, and (iv) mixed precision in selected portions of the code. A series of performance benchmarks has been carried out on leadership high-performance computing systems, showing a substantial speedup of the GPU-accelerated version of WEST with respect to its CPU version. Good strong and weak scaling is demonstrated using up to 25 920 GPUs. Finally, we showcase the capability of the GPU version of WEST for large-scale, full-frequency GW calculations of realistic systems, e.g., a nanostructure, an interface, and a defect, comprising up to 10 368 valence electrons.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On the optical anisotropy in 2D metal-halide perovskites

We develop a better understanding of the many contributing factors that give rise to extreme optical anisotropy in 2D perovskites, and we then show that spin-coated films can exhibit excellent order comparable with exfoliated crystals.

14 SOLAR ENERGY↗

Accurate frozen core approximation for all-electron density-functional theory

We implement and benchmark the frozen core approximation, a technique commonly adopted in electronic structure theory to reduce the computational cost by means of mathematically fixing the chemically inactive core electron states. The accuracy and efficiency of this approach are well controlled by a single parameter, the number of frozen orbitals. Explicit corrections for the frozen core orbitals and the unfrozen valence orbitals are introduced, safeguarding against seemingly minor numerical deviations from the assumed orthonormality conditions of the basis functions. A speedup of over twofold can be achieved for the diagonalization step in all-electron density-functional theory simulations containing heavy elements, without any accuracy degradation in terms of the electron density, total energy, and atomic forces. This is demonstrated in a benchmark study covering 103 materials across the Periodic Table and a large-scale simulation of CsPbBr3 with 2560 atoms. Our study provides a rigorous benchmark of the precision of the frozen core approximation (sub-meV per atom for frozen core orbitals below −200 eV) for a wide range of test cases and for chemical elements ranging from Li to Po. The algorithms discussed here are implemented in the open-source Electronic Structure Infrastructure software package.

Yu, Victor Wen-zhe↗

GPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and hermitian eigenproblems

The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. Furthermore, we benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU–GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.

97 MATHEMATICS AND COMPUTING↗

ELSI — An open infrastructure for electronic structure solvers

Routine applications of electronic structure theory to molecules and periodic systems need to compute the electron density from given Hamiltonian and, in case of non-orthogonal basis sets, overlap matrices. System sizes can range from few to thousands or, in some examples, millions of atoms. Different discretization schemes (basis sets) and different system geometries (finite non-periodic vs. infinite periodic boundary conditions) yield matrices with different structures. The ELectronic Structure Infrastructure (ELSI) project provides an open-source software interface to facilitate the implementation and optimal use of high-performance solver libraries covering cubic scaling eigensolvers, linear scaling density-matrix-based algorithms, and other reduced scaling methods in between. In this paper, we present recent improvements and developments inside ELSI, mainly covering (1) new solvers connected to the interface, (2) matrix layout and communication adapted for parallel calculations of periodic and/or spin-polarized systems, (3) routines for density matrix extrapolation in geometry optimization and molecular dynamics calculations, and (4) general utilities such as parallel matrix I/O and JSON output. The ELSI interface has been integrated into four electronic structure code projects (DFTB+, DGDFT, FHI-aims, SIESTA), allowing us to rigorously benchmark the performance of the solvers on an equal footing. Based on results of a systematic set of large-scale benchmarks performed with Kohn–Sham density-functional theory and density-functional tight-binding theory, we identify factors that strongly affect the efficiency of the solvers, and propose a decision layer that assists with the solver selection process. As a result, we describe a reverse communication interface encoding matrix-free iterative solver strategies that are amenable, e.g., for use with planewave basis sets.

97 MATHEMATICS AND COMPUTING↗

GPU acceleration of all-electron electronic structure theory using localized numeric atom-centered basis functions

We present an implementation of all-electron density-functional theory for massively parallel GPU-based platforms, using localized atom-centered basis functions and real-space integration grids. Special attention is paid to domain decomposition of the problem on non-uniform grids, which enables compute- and memory-parallel execution across thousands of nodes for real-space operations, e.g. the update of the electron density, the integration of the real-space Hamiltonian matrix, and calculation of Pulay forces. To assess the performance of our GPU implementation, we performed benchmarks on three different architectures using a 103-material test set. We find that operations which rely on dense serial linear algebra show dramatic speedups from GPU acceleration: in particular, SCF iterations including force and stress calculations exhibit speedups ranging from 4.5 to 6.6. For the architectures and problem types investigated here, this translates to an expected overall speedup between 3–4 for the entire calculation (including non-GPU accelerated parts), for problems featuring several tens to hundreds of atoms. Additional calculations for a 375-atom Bi2Se3 bilayer show that the present GPU strategy scales for large-scale distributed-parallel simulations.

42 ENGINEERING↗