Search NASASearch

Engineering topics

Feiguin, Adrian

Publications and source records attributed to Feiguin, Adrian.

3D Heisenberg universality in the van der Waals antiferromagnet NiPS 3

Van der Waals (vdW) magnetic materials are comprised of layers of atomically thin sheets, making them ideal platforms for studying magnetism at the two-dimensional (2D) limit. These materials are at the center of a host of novel types of experiments, however, there are notably few pathways to directly probe their magnetic structure. We confirm the magnetic order within a single crystal of NiPS 3 and show it can be accessed with resonant elastic X-ray diffraction along the edge of the vdW planes in a carefully grown crystal by detecting structurally forbidden resonant magnetic X-ray scattering. We find the magnetic order parameter has a critical exponent of β ~ 0.36, indicating that the magnetism of these vdW crystals is more adequately characterized by the three-dimensional (3D) Heisenberg universality class. We verify these findings with first-principles density functional theory, Monte-Carlo simulations, and density matrix renormalization group calculations.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING