Search NASA⌕ Search

SEARCH · Search NASA

Results for “Tensor rank determination”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

General-Purpose Bayesian Tensor Learning With Automatic Rank Determination and Uncertainty Quantification

A major challenge in many machine learning tasks is that the model expressive power depends on model size. Low-rank tensor methods are an efficient tool for handling the curse of dimensionality in many large-scale machine learning models. The major challenges in training a tensor learning model include how to process the high-volume data, how to determine the tensor rank automatically, and how to estimate the uncertainty of the results. While existing tensor learning focuses on a specific task, this paper proposes a generic Bayesian framework that can be employed to solve a broad class of tensor learning problems such as tensor completion, tensor regression, and tensorized neural networks. We develop a low-rank tensor prior for automatic rank determination in nonlinear problems. Our method is implemented with both stochastic gradient Hamiltonian Monte Carlo (SGHMC) and Stein Variational Gradient Descent (SVGD). We compare the automatic rank determination and uncertainty quantification of these two solvers. We demonstrate that our proposed method can determine the tensor rank automatically and can quantify the uncertainty of the obtained results. We validate our framework on tensor completion tasks and tensorized neural network training tasks.

Bayesian inference↗

Towards Compact Neural Networks via End-to-End Training: A Bayesian Tensor Approach with Automatic Rank Determination

Post-training model compression can reduce the inference costs of deep neural networks, but uncompressed training still consumes enormous hardware resources and energy. To enable low-energy training on edge devices, it is highly desirable to directly train a compact neural network from scratch with a low memory cost. Low-rank tensor decomposition is an effective approach to reduce the memory and computing costs of large neural networks. However, directly training low-rank tensorized neural networks is a very challenging task because it is hard to determine a proper tensor rank a priori, and the tensor rank controls both model complexity and accuracy. Here, this paper presents a novel end-to-end framework for low-rank tensorized training. We first develop a Bayesian model that supports various low-rank tensor formats (e.g., CANDECOMP/PARAFAC, Tucker, tensor-train, and tensor-train matrix) and reduces neural network parameters with automatic rank determination during training. Then we develop a customized Bayesian solver to train large-scale tensorized neural networks. Our training methods shows orders-of-magnitude parameter reduction and little accuracy loss (or even better accuracy) in the experiments. On a very large deep learning recommendation system with over 4.2 ×10 9 model parameters, our method can reduce the parameter number to 1.6 ×10 5 automatically in the training process (i.e., by 2.6 ×10 4 times) while achieving almost the same accuracy. Code is available at https://github.com/colehawkins/bayesian-tensor-rank-determination.

compact neural networks↗

Bayesian tensorized neural networks with automatic rank selection

Tensor decomposition is an effective approach to compress over-parameterized neural networks and to enable their deployment on resource-constrained hardware platforms. However, directly applying tensor compression in the training process is a challenging task due to the difficulty of choosing a proper tensor rank. In order to address this challenge, this paper proposes a low-rank Bayesian tensorized neural network. Our Bayesian method performs automatic model compression via an adaptive tensor rank determination. We also present approaches for posterior density calculation and maximum a posteriori (MAP) estimation for the end-to-end training of our tensorized neural network. Here, we provide experimental validation on a two-layer fully connected neural network, a 6-layer CNN and a 110-layer residual neural network where our work produces 7.4x to 137x more compact neural networks directly from the training while achieving high prediction accuracy.

Low-rank tensor↗

Anisotropic Cahn-Hilliard free energy and interfacial energies for binary alloys with pairwise interactions

The original Cahn-Hilliard derivation of the contribution of compositional inhomogeneity to the free energy of a binary alloy with pairwise interactions is extended to include higher-order inhomogeneity terms. For alloys on a cubic lattice, the coefficient of the first inhomogeneity is a second-rank tensor and reduces to a scalar, but it is shown that the second order and the third order inhomogeneity terms are weighted by fourth-rank and sixth-rank tensors, thus resulting in anisotropic contributions. Furthermore, each interaction shell generates a unique set of inhomogeneity coefficients that is determined by the set of vectors connecting an atom to its neighbors on that shell. These coefficients are calculated for fcc and bcc alloys with interactions up fourth nearest neighbors. Phase field simulations based on these extended Cahn-Hilliard free energies are performed to measure interface free energies along specific crystallographic directions as a function of temperature, and to obtain the equilibrium shape of precipitates. Interface free energies, and the resulting anisotropies, are compared to those obtained by discrete models and Monte Carlo simulations.

36 MATERIALS SCIENCE↗

Zero-truncated Poisson regression for sparse multiway count data corrupted by false zeros

Abstract We propose a novel statistical inference methodology for multiway count data that is corrupted by false zeros that are indistinguishable from true zero counts. Our approach consists of zero-truncating the Poisson distribution to neglect all zero values. This simple truncated approach dispenses with the need to distinguish between true and false zero counts and reduces the amount of data to be processed. Inference is accomplished via tensor completion that imposes low-rank tensor structure on the Poisson parameter space. Our main result shows that an $N$-way rank-$R$ parametric tensor $\boldsymbol{\mathscr{M}}\in (0,\infty )^{I\times \cdots \times I}$ generating Poisson observations can be accurately estimated by zero-truncated Poisson regression from approximately $IR^2\log _2^2(I)$ non-zero counts under the nonnegative canonical polyadic decomposition. Our result also quantifies the error made by zero-truncating the Poisson distribution when the parameter is uniformly bounded from below. Therefore, under a low-rank multiparameter model, we propose an implementable approach guaranteed to achieve accurate regression in under-determined scenarios with substantial corruption by false zeros. Several numerical experiments are presented to explore the theoretical results.

97 MATHEMATICS AND COMPUTING↗

PyMTensor 1.0.0

SAND2021-1654 O The Python Material Tensor code (PyMTensor) applies crystal symmetries to arbitrary material tensors. It uses symbolic algebra to exactly determine the zero tensor components and the relationships between the nonzero components for all crystallographic point groups. It is capable of computing components of higher-order material tensors needed when treating nonlinear effects. It has a simple and flexible interface for inputting tensor rank information and any known symmetries between tensor indices.Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Jensen, DanielS↗

6D large charge and 2D Virasoro blocks

We compute observables in the interacting rank-one 6D 𝒩 =(2,0) superconformal field theory (SCFT) at large 𝑅-charge. We focus on correlators involving Φ 𝑛 , namely symmetric products of the bottom component of the supermultiplet containing the stress tensor. By using the moduli space effective action and methods from the large-charge expansion, we compute the operator product expansion coefficients ⟨Φ 𝑛 ⁢Φ 𝑚 ⁢Φ 𝑛+𝑚 ⟩ in an expansion in 1/𝑛. The coefficients of the expansion are only partially determined from the 6D perspective, but we manage to fix them order-by-order in 1/𝑛 numerically by utilizing the 6⁢D/2⁢D correspondence. This is made possible by the fact that this 6D observable can be extracted in 2D from a specific double-scaling limit of the vacuum Virasoro block, which can be efficiently computed numerically. We also extend the computation to higher-rank SCFTs, and discuss various applications of our results to 6D as well as 2D.

classical solutions in field theory↗

Minimal parameter solution of the orthogonal matrix differential equation

As demonstrated in this work, all orthogonal matrices solve a first order differential equation. The straightforward solution of this equation requires n sup 2 integrations to obtain the element of the nth order matrix. There are, however, only n(n-1)/2 independent parameters which determine an orthogonal matrix. The questions of choosing them, finding their differential equation and expressing the orthogonal matrix in terms of these parameters are considered. Several possibilities which are based on attitude determination in three dimensions are examined. It is shown that not all 3-D methods have useful extensions to higher dimensions. It is also shown why the rate of change of the matrix elements, which are the elements of the angular rate vector in 3-D, are the elements of a tensor of the second rank (dyadic) in spaces other than three dimensional. It is proven that the 3-D Gibbs vector (or Cayley Parameters) are extendable to other dimensions. An algorithm is developed employing the resulting parameters, which are termed Extended Rodrigues Parameters, and numerical results are presented of the application of the algorithm to a fourth order matrix.

Bar-Itzhack, Itzhack Y.↗

Minimal parameter solution of the orthogonal matrix differential equation

As demonstrated in this work, all orthogonal matrices solve a first order differential equation. The straightforward solution of this equation requires n sup 2 integrations to obtain the element of the nth order matrix. There are, however, only n(n-1)/2 independent parameters which determine an orthogonal matrix. The questions of choosing them, finding their differential equation and expressing the orthogonal matrix in terms of these parameters are considered. Several possibilities which are based on attitude determination in three dimensions are examined. It is shown that not all 3-D methods have useful extensions to higher dimensions. It is also shown why the rate of change of the matrix elements, which are the elements of the angular rate vector in 3-D, are the elements of a tensor of the second rank (dyadic) in spaces other than three dimensional. It is proven that the 3-D Gibbs vector (or Cayley Parameters) are extendable to other dimensions. An algorithm is developed employing the resulting parameters, which are termed Extended Rodrigues Parameters, and numerical results are presented of the application of the algorithm to a fourth order matrix.

Baritzhack, Itzhack Y.↗

Minimal parameter solution of the orthogonal matrix differential equation

As demonstrated in this work, all orthogonal matrices solve a first order differential equation. The straightforward solution of this equation requires n sup 2 integrations to obtain the element of the nth order matrix. There are, however, only n(n-1)/2 independent parameters which determine an orthogonal matrix. The questions of choosing them, finding their differential equation and expressing the orthogonal matrix in terms of these parameters are considered. Several possibilities which are based on attitude determination in three dimensions are examined. It is shown that not all 3-D methods have useful extensions to higher dimensions. It is also shown why the rate of change of the matrix elements, which are the elements of the angular rate vector in 3-D, are the elements of a tensor of the second rank (dyadic) in spaces other than three dimensional. It is proven that the 3-D Gibbs vector (or Cayley Parameters) are extendable to other dimensions. An algorithm is developed emplying the resulting parameters, which are termed Extended Rodrigues Parameters, and numerical results are presented of the application of the algorithm to a fourth order matrix.

Bar-Itzhack, Itzhack Y.↗

Tensor Decompositions for Count Data that Leverage Stochastic and Deterministic Optimization

There is growing interest to extend low-rank matrix decompositions to multi-way arrays, or tensors. One fundamental low-rank tensor decomposition is the canonical polyadic decomposition (CPD). The challenge of fitting a low-rank, nonnegative CPD model to Poisson-distributed count data is of particular interest. Several popular algorithms use local search methods to approximate the global maximum likelihood estimator from local minima. Simultaneously, a recent trend in theoretical computer science and numerical linear algebra leverages randomization to solve very large, hard problems. The typical approach is to use randomization for a fast approximation and determinism for refinement to yield effective algorithms with theoretical guarantees. Two popular algorithms for Poisson CPD reflect that emergent dichotomy: CP Alternating Poisson Regression is a deterministic algorithm and Generalized Canonical Polyadic decomposition makes use of stochastic algorithms in several variants. This work extends recent work to develop two new methods that leverage randomized and deterministic algorithms for improved accuracy and performance.

97 MATHEMATICS AND COMPUTING↗

General field evaluation in high-order meshes on GPUs

Robust and scalable function evaluation at any arbitrary point in the finite/spectral element mesh is required for querying the partial differential equation solution at points of interest, comparison of solution between different meshes, and Lagrangian particle tracking. This is a challenging problem, particularly for high-order unstructured meshes partitioned in parallel with MPI, as it requires identifying the element that overlaps a given point and computing the corresponding reference space coordinates. Here, we present a robust and efficient technique for general field evaluation in large-scale high-order meshes with quadrilaterals and hexahedra. In the proposed method, a combination of globally partitioned and processor-local maps are used to first determine a list of candidate MPI ranks, and then locally candidate elements that could contain a given point. Next, element-wise bounding boxes further reduce the list of candidate elements. Finally, Newton’s method with trust region is used to determine the overlapping element and corresponding reference space coordinates. Since GPU-based architectures have become popular for accelerating computational analyses using meshes with tensor-product elements, specialized kernels have been developed to utilize the proposed methodology on GPUs. The method is also extended to enable general field evaluation on surface meshes. The paper concludes by demonstrating the use of the proposed method in various applications ranging from mesh-to-mesh transfer during r-adaptivity to Lagrangian particle tracking.

97 MATHEMATICS AND COMPUTING↗

GPU acceleration of rank-reduced coupled-cluster singles and doubles

Here, we have developed a graphical processing unit (GPU) accelerated implementation of our recently introduced rank-reduced coupled-cluster singles and doubles (RR-CCSD) method. RR-CCSD introduces a low-rank approximation of the doubles amplitudes. This is combined with a low-rank approximation of the electron repulsion integrals via Cholesky decomposition. The result of these two low-rank approximations is the replacement of the usual fourth-order CCSD tensors with products of second- and third-order tensors. In our implementation, only a single fourth-order tensor must be constructed as an intermediate during the solution of the amplitude equations. Owing in large part to the compression of the doubles amplitudes, the GPU-accelerated implementation shows excellent parallel efficiency (95% on eight GPUs). Our implementation can solve the RR-CCSD equations for up to 400 electrons and 1550 basis functions—roughly 50% larger than the largest canonical CCSD computations that have been performed on any hardware. In addition to increased scalability, the RR-CCSD computations are faster than the corresponding CCSD computations for all but the smallest molecules. We test the accuracy of RR-CCSD for a variety of chemical systems including up to 1000 basis functions and determine that accuracy to better than 0.1% error in the correlation energy can be achieved with roughly 95% compression of the ov space for the largest systems considered. We also demonstrate that conformational energies can be predicted to be within 0.1 kcal mol -1 with efficient compression applied to the wavefunction. Finally, we find that low-rank approximations of the CCSD doubles amplitudes used in the similarity transformation of the Hamiltonian prior to a conventional equation-of-motion CCSD computation will not introduce significant errors (on the order of a few hundredths of an electronvolt) into the resulting excitation energies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Embedding of attitude determination in n-dimensional spaces

The problem of attitude determination in n-dimensional spaces is addressed. The proper parameters are found, and it is shown that not all three-dimensional methods have useful extensions to higher dimensions. It is demonstrated that Rodriguez parameters are conveniently extendable to other dimensions. An algorithm for using these parameters in the general n-dimensional case is developed and tested with a four-dimensional example. The correct mathematical description of angular velocities is addressed, showing that angular velocity in n dimensions cannot be represented by a vector but rather by a tensor of the second rank. Only in three dimensions can the angular velocity be described by a vector.

Bar-Itzhack, Itzhack Y.↗

The Effective Fragment Molecular Orbital Method: Achieving High Scalability and Accuracy for Large Systems

The effective fragment molecular orbital (EFMO) method has been developed to predict the total energy of a very large molecular system accurately (with respect to the underlying quantum mechanical method) and efficiently by taking advantage of the locality of strong chemical interactions and employing a two-level hierarchical parallelism. The accuracy of the EFMO method is partly attributed to the accurate and robust intermolecular interaction prediction between distant fragments, in particular, the many-body polarization and dispersion effects, which require the generation of static and dynamic polarizability tensors by solving the coupled perturbed Hartree–Fock (CPHF) and time-dependent HF (TDHF) equations, respectively. Solving the CPHF and TDHF equations is the main EFMO computational bottleneck due to the inefficient (serial) and I/O-intensive implementation of the CPHF and TDHF solvers. In this work, the efficiency and scalability of the EFMO method are significantly improved with a new CPU memory-based implementation for solving the CPHF and TDHF equations that are parallelized by either message passing interface (MPI) or hybrid MPI/OpenMP. Here, the accuracy of the EFMO method is demonstrated for both covalently bonded systems and noncovalently bound molecular clusters by systematically examining the effects of basis sets and a key distance-related cutoff parameter, R cut . R cut determines whether a fragment pair (dimer) is treated by the chosen ab initio method or calculated using the effective fragment potential (EFP) method (separated dimers). Decreasing the value of Rcut increases the number of separated (EFP) dimers, thereby decreasing the computational effort. It is demonstrated that excellent accuracy (<1 kcal/mol error per fragment) can be achieved when using a sufficiently large basis set with diffuse functions coupled with a small R cut value. With the new parallel implementation, the total EFMO wall time is substantially reduced, especially with a high number of MPI ranks. Given a sufficient workload, nearly ideal strong scaling is achieved for the CPHF and TDHF parts of the calculation. For the first time, EFMO calculations with the inclusion of long-range polarization and dispersion interactions on a hydrated mesoporous silica nanoparticle with explicit water solvent molecules (more than 15k atoms) are achieved on a massively parallel supercomputer using nearly 1000 physical nodes. In addition, EFMO calculations on the carbinolamine formation step of an amine-catalyzed aldol reaction at the nanoscale with explicit solvent effects are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗