Search NASA⌕ Search

SEARCH · Search NASA

Results for “matrix decomposition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Identifying Entangled Physics Relationships Through Sparse Matrix Decomposition to Inform Plasma Fusion Design

We report a sustainable burn platform through inertial confinement fusion (ICF) has been an ongoing challenge for over 50 years. Mitigating engineering limitations and improving the current design involves an understanding of the complex coupling of physical processes. While sophisticated simulation codes are used to model ICF implosions, these tools contain necessary numerical approximation but miss physical processes that limit predictive capability. Identification of relationships between controllable design inputs to ICF experiments and measurable outcomes (e.g., neutron yield, neutron velocity, areal density) from performed experiments can help guide the future design of experiments and development of simulation codes, to potentially improve the accuracy of the computational models used to simulate ICF experiments. We use sparse matrix decomposition methods to identify clusters of a few related design variables. Sparse principal component analysis (SPCA) identifies groupings that are related to the physical origin of the variables (laser, hohlraum, and capsule). A variable importance analysis finds that in addition to variables highly correlated with neutron yield, such as picket power and laser energy, variables that represent a dramatic change of the ICF design, such as number of pulse steps, are also very important. The obtained sparse components are then used to train a random forest (RF) regression surrogate for predicting total yield. The RF performance on the training and testing data compares with the performance of the RF trained using all the design variables considered. This work is intended to inform design changes in future ICF experiments by augmenting the expert intuition and simulation results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

MatRIS: Addressing the Challenges for Portability and Heterogeneity Using Tasking for Matrix Decomposition (Cholesky)

The ubiquitous in-node heterogeneity of HPC and cloud computing platforms makes software portability and performance optimization extremely challenging. Described here, the MatRIS multilevel math library abstraction framework employs tasking to alleviate these difficulties. MatRIS includes the IRIS task-based runtime on the bottom level and exposes different layers of abstraction to render algorithms architecturally agnostic. MatRIS ensures the decomposition and creation of tasks that represent the necessary encapsulation of the optimized kernels from both vendor and open-source math libraries. Once built, MatRIS can select different combinations of accelerators at runtime, making it portable even on diverse heterogeneous architectures. By leveraging the IRIS runtime’s features for managing heterogeneity, MatRIS deploys algorithms that remove the need to specify orchestration and data transfer. This study describes how the serial task abstraction of a tiled Cholesky factorization is made portable and scalable in the case of multi-device and multi-vendor heterogeneity on a node with NVIDIA and AMD GPUs by using MatRIS. First, we demonstrate that Cholesky in MatRIS provides multi-GPU scalability that offers competitive performance versus cuSolverMG. Then, we present the challenges and opportunities for heterogeneous execution.

Monil, M. A. H.↗

Identifying Entangled Physics Relationships through Sparse Matrix Decomposition to Inform Plasma Fusion Design [Slides]

The National Ignition Facility (NIF), is a large laser-based inertial confinement fusion (ICF) research device and various input variables in the experimental data are described. The overview included sections on: high-dimensional experimental dataset; surrogate model selection; ML to interpret complex coupling between inputs; importance of variables; Surrogate performance; and, future work.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Compression of tokamak boundary plasma simulation data using a maximum volume algorithm for matrix skeleton decomposition

This report demonstrates satisfactory data compression of SOLPS-ITER simulation output ranging from 2D fields, 1D profiles, and 0D scalar variables with a novel matrix decomposition approach. The singular value decomposition (SVD) scales poorly for large matrix sizes and is unsuited to the application on high dimensional data common to fusion plasma physics simulation. In this work, we employ the columns-submatrix-rows (CUR) matrix factorization technique in order to compute a low-rank approximation up to two orders of magnitude faster than the SVD, but within a nominal L2-norm relative error of ε = 10 –2 . In addition, the CUR approach maintains the original format of the data, in its extracted columns and rows, allowing for interpretable data storage at the original resolution of the simulation. We utilize an iterative algorithm to compute the CUR decomposition of simulation output by maximizing the volume, or linearly independent information content, of a low-rank submatrix contained within the data. Experiments over $\textit{n} × \textit{n}$ randomized test matrices with embedded rank-deficient features show that this maximum volume implementation of CUR matrix approximation has reduced asymptotic computational complexity on the order of n compared to the SVD, which scales approximately as $n^3$. These results show that the CUR technique can be used to effectively select time step snapshots (columns) of over 140 SOLPS-ITER output variables and the associated discretized coordinate timeseries (rows) allowing for reconstruction of the complete simulation dynamics.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Quantum Fourier transform revisited

Summary The fast Fourier transform (FFT) is one of the most successful numerical algorithms of the 20th century and has found numerous applications in many branches of computational science and engineering. The FFT algorithm can be derived from a particular matrix decomposition of the discrete Fourier transform (DFT) matrix. In this paper, we show that the quantum Fourier transform (QFT) can be derived by further decomposing the diagonal factors of the FFT matrix decomposition into products of matrices with Kronecker product structure. We analyze the implication of this Kronecker product structure on the discrete Fourier transform of rank‐1 tensors on a classical computer. We also explain why such a structure can take advantage of an important quantum computer feature that enables the QFT algorithm to attain an exponential speedup on a quantum computer over the FFT algorithm on a classical computer. Further, the connection between the matrix decomposition of the DFT matrix and a quantum circuit is made. We also discuss a natural extension of a radix‐2 QFT decomposition to a radix‐ d QFT decomposition. No prior knowledge of quantum computing is required to understand what is presented in this paper. Yet, we believe this paper may help readers to gain some rudimentary understanding of the nature of quantum computing from a matrix computation point of view.

Camps, Daan↗

Porting fragmentation methods to GPUs using an OpenMP API: Offloading the resolution-of-the-identity second-order Møller–Plesset perturbation method

Here, using an OpenMP Application Programming Interface, the resolution-of-the-identity second-order Møller–Plesset perturbation (RI-MP2) method has been off-loaded onto graphical processing units (GPUs), both as a standalone method in the GAMESS electronic structure program and as an electron correlation energy component in the effective fragment molecular orbital (EFMO) framework. First, a new scheme has been proposed to maximize data digestion on GPUs that subsequently linearizes data transfer from central processing units (CPUs) to GPUs. Second, the GAMESS Fortran code has been interfaced with GPU numerical libraries (e.g., NVIDIA cuBLAS and cuSOLVER) for efficient matrix operations (e.g., matrix multiplication, matrix decomposition, and matrix inversion). The standalone GPU RI-MP2 code shows an increasing speedup of up to 7.5× using one NVIDIA V100 GPU with one IBM 42-core P9 CPU for calculations on fullerenes of increasing size from 40 to 260 carbon atoms using the 6-31G(d)/cc-pVDZ-RI basis sets. A single Summit node with six V100s can compute the RI-MP2 correlation energy of a cluster of 175 water molecules using the correlation consistent basis sets cc-pVDZ/cc-pVDZ-RI containing 4375 atomic orbitals and 14 700 auxiliary basis functions in ~0.85 h. In the EFMO framework, the GPU RI-MP2 component shows near linear scaling for a large number of V100s when computing the energy of an 1800-atom mesoporous silica nanoparticle in a bath of 4000 water molecules. The parallel efficiencies of the GPU RI-MP2 component with 2304 and 4608 V100s are 98.0% and 96.1%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Pass-efficient methods for compression of high-dimensional turbulent flow data

The future of high-performance computing, specifically on future Exascale computers, will presumably see memory capacity and bandwidth fail to keep pace with data generated, for instance, from massively parallel partial differential equation (PDE) systems. Current strategies proposed to address this bottleneck entail the omission of large fractions of data, as well as the incorporation of in situ compression algorithms to avoid overuse of memory. To ensure that post-processing operations are successful, this must be done in a way that a sufficiently accurate representation of the solution is stored. Moreover, in situations where the input/output system becomes a bottleneck in analysis, visualization, etc., or the execution of the PDE solver is expensive, the number of passes made over the data must be minimized. In the interest of addressing this problem, this work focuses on the utility of pass-efficient, parallelizable, low-rank, matrix decomposition methods in compressing high-dimensional simulation data from turbulent flows. Additionally, a particular emphasis is placed on using coarse representation of the data – compatible with the PDE discretization grid – to accelerate the construction of the low-rank factorization. This includes the presentation of a novel single-pass matrix decomposition algorithm for computing the so-called interpolative decomposition. The methods are described extensively and numerical experiments on two turbulent channel flow data are performed. In the first (unladen) channel flow case, compression factors exceeding 400 are achieved while maintaining accuracy with respect to first- and second-order flow statistics. In the particle-laden case, compression factors of 100 are achieved and the compressed data is used to recover particle velocities. These results show that these compression methods can enable efficient computation of various quantities of interest in both the carrier and disperse phases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Tensor decompositions for count data that leverage stochastic and deterministic optimization

There is growing interest to extend low-rank matrix decompositions to multi-way arrays, or tensors. One fundamental low-rank tensor decomposition is the canonical polyadic decomposition (CPD). The challenge of fitting a low-rank, nonnegative CPD model to Poisson-distributed count data is of particular interest. Several popular algorithms use local search methods to approximate the maximum likelihood estimator (MLE) of the Poisson CPD model. Here, this work presents two new algorithms that extend state-of-the-art local methods for Poisson CPD. Hybrid GCP-CPAPR combines Generalized Canonical Decomposition (GCP) with stochastic optimization and CP Alternating Poisson Regression (CPAPR), a deterministic algorithm, to increase the probability of converging to the MLE over either method used alone. Restarted CPAPR with SVDrop uses a heuristic based on the singular values of the CPD model unfoldings to identify convergence toward optimizers that are not the MLE and restarts within the feasible domain of the optimization problem, thus reducing overall computational cost when using a multi-start strategy. We provide empirical evidence that indicates our approaches outperform existing methods with respect to converging to the Poisson CPD MLE.

CPAPR↗

Tensor Decompositions for Count Data that Leverage Stochastic and Deterministic Optimization

There is growing interest to extend low-rank matrix decompositions to multi-way arrays, or tensors. One fundamental low-rank tensor decomposition is the canonical polyadic decomposition (CPD). The challenge of fitting a low-rank, nonnegative CPD model to Poisson-distributed count data is of particular interest. Several popular algorithms use local search methods to approximate the global maximum likelihood estimator from local minima. Simultaneously, a recent trend in theoretical computer science and numerical linear algebra leverages randomization to solve very large, hard problems. The typical approach is to use randomization for a fast approximation and determinism for refinement to yield effective algorithms with theoretical guarantees. Two popular algorithms for Poisson CPD reflect that emergent dichotomy: CP Alternating Poisson Regression is a deterministic algorithm and Generalized Canonical Polyadic decomposition makes use of stochastic algorithms in several variants. This work extends recent work to develop two new methods that leverage randomized and deterministic algorithms for improved accuracy and performance.

97 MATHEMATICS AND COMPUTING↗

Randomized Projection for Rank-Revealing Matrix Factorizations and Low-Rank Approximations

Rank-revealing matrix decompositions provide an essential tool in spectral analysis of matrices, including the Singular Value Decomposition (SVD) and related low-rank approximation techniques. QR with Column Pivoting (QRCP) is usually suitable for these purposes, but it can be much slower than the unpivoted QR algorithm. For large matrices, the difference in performance is due to increased communication between the processor and slow memory, which QRCP needs in order to choose pivots during decomposition. Our main algorithm, Randomized QR with Column Pivoting (RQRCP), uses randomized projection to make pivot decisions from a much smaller sample matrix, which we can construct to reside in a faster level of memory than the original matrix. This technique may be understood as trading vastly reduced communication for a controlled increase in uncertainty during the decision process. Furthermore, for rank-revealing purposes, the selection mechanism in RQRCP produces results that are the same quality as the standard algorithm, but with performance near that of unpivoted QR (often an order of magnitude faster for large matrices). Additionally, we also propose two formulas that facilitate further performance improvements. The first efficiently updates sample matrices to avoid computing new randomized projections. The second avoids large trailing updates during the decomposition in truncated low-rank approximations. Our truncated version of RQRCP also provides a key initial step in our truncated SVD approximation, TUXV. These advances open up a new performance domain for large matrix factorizations that will support efficient problem-solving techniques for challenging applications in science, engineering, and data analysis.

97 MATHEMATICS AND COMPUTING↗

High-dimensional data analytics in civil engineering: A review on matrix and tensor decomposition

Recent developments in sensing and monitoring techniques have led to the generation of high-dimensional data in the field of civil engineering. High-dimensional data analytics methods have thus been developed to interpret such complex data. Among the different high-dimensional data analytics techniques, matrix and tensor decomposition methods have acquired a notable interest in the civil engineering community over the past decade. Due to their unique ability to deal with highly redundant and correlated data, these methods are establishing themselves as promising and efficient tools to analyze high-dimensional data in the civil engineering arena. In this paper, high-dimensional data is referred to as a data set in which the number of features is comparable or larger than the number of observations. This review paper aims to summarize the applications of matrix and tensor decomposition methods in civil engineering over the last decade. The survey begins with a general overview of matrix and tensor decomposition followed by highlighting their significance in the field. Afterward, various applications of these high-dimensional data analytics methods in civil engineering are presented, while the advantages offered by these methods are discussed. Lastly, challenges and potential research avenues for employing matrix and tensor decomposition and future emerging trends for their novel use are highlighted.

42 ENGINEERING↗

Time-dependent-bases with local CUR decomposition method for accelerating turbulent combustion simulations

Here, this study presents a novel reduced-order modeling framework, Time-Dependent Bases with Local CUR decomposition (TDB-L-CUR), designed to efficiently and accurately approximate the species transport equations in reacting flow simulations. The method extends the existing TDB-CUR approach for chemically reacting flows (Jung et al. Comput. Methods Appl. Mech. Engrg. 437 (2025) 117758), which leverages matrix decomposition techniques to form a global-in-space, time-dependent low-dimensional manifold. While TDB-CUR performs well in homogeneous systems, it may be less well-suited to spatially heterogeneous systems such as turbulent flames, where higher-rank approximations are typically required. The proposed TDB-L-CUR framework introduces two methodological extensions to the baseline approach. First, it applies unsupervised clustering to partition the physical domain into distinct regions, enabling spatially localized manifold construction, thereby reducing the rank required for the reduced-order representation. Second, it incorporates a computational singular perturbation (CSP)-based scheme for identifying and penalizing fast species, allowing for spatio-temporally adaptive mitigation of chemical stiffness. The proposed framework is validated on a hierarchy of test cases, including a one-dimensional premixed flame, a two-dimensional nonpremixed ignition case with vortex interaction, and a three-dimensional turbulent premixed flame. TDB-L-CUR significantly improves accuracy over TDB-CUR while further reducing computational cost. The fully on-the-fly formulation of TDB-L-CUR (i.e., requiring no offline training or prior knowledge) makes it a robust and scalable tool for reduced-order modeling of reactive flows.

Local manifold↗

A dictionary learning algorithm for compression and reconstruction of streaming data in preset order

There has been an emerging interest in developing and applying dictionary learning (DL) to process massive datasets in the last decade. Many of these efforts, however, focus on employing DL to compress and extract a set of important features from data, while considering restoring the original data from this set a secondary goal. On the other hand, although several methods are able to process streaming data by updating the dictionary incrementally as new snapshots pass by, most of those algorithms are designed for the setting where the snapshots are randomly drawn from a probability distribution. In this paper, we present a new DL approach to compress and denoise massive dataset in real time, in which the data are streamed through in a preset order (instances are videos and temporal experimental data), so at any time, we can only observe a biased sample set of the whole data. Here, our approach incrementally builds up the dictionary in a relatively simple manner: if the new snapshot is adequately explained by the current dictionary, we perform a sparse coding to find its sparse representation; otherwise, we add the new snapshot to the dictionary, with a Gram-Schmidt process to maintain the orthogonality. To compress and denoise noisy datasets, we apply the denoising to the snapshot directly before sparse coding, which deviates from traditional dictionary learning approach that achieves denoising via sparse coding. Compared to full-batch matrix decomposition methods, where the whole data is kept in memory, and other mini-batch approaches, where unbiased sampling is often assumed, our approach has minimal requirement in data sampling and storage: i) each snapshot is only seen once then discarded, and ii) the snapshots are drawn in a preset order, so can be highly biased. Through experiments on climate simulations and scanning transmission electron microscopy (STEM) data, we demonstrate that the proposed approach performs competitively to those methods in data reconstruction and denoising.

97 MATHEMATICS AND COMPUTING↗

Segmenting the target audience for transportation demand management programs: An investigation between mode shift and individual characteristics

Understanding the characteristics of travelers and situations that are more likely to switch from single occupancy vehicles (SOV) to more sustainable mobility options can help transportation agencies develop better behavior change strategies. Existing mode shift programs, however, rarely identify their target audience, which could make the programs less effective. Here, to investigate the demographics and trip characteristics of those who are more likely to shift from SOV to carpooling, public transit, and micromobility (e.g. walking, biking), a matrix decomposition audience segmentation method is proposed and applied on travel survey data. The results of this research show that for short-distance leisure or social trips, young and middle-income people are more likely to shift to these mobility options from SOV. This study provides a comprehensive and profound understanding of the characteristics of the target audience that can inform policies to create more targeted behavior change strategies to reach their sustainability goals.

42 ENGINEERING↗

Fast calculation of the tokamak vertical instability

There has been recent interest in fast calculations of the tokamak axisymmetric vertical instability for real time feedback control purposes. It is shown that the maximum eigenvalue for the basic rigid version of this stability problem can be obtained by finding the positive root to a simple scalar function. This function can be generalized to include plasma mass and has complexity linear in the number of conductive elements. The formulation is based on standard matrix decompositions of the fixed-geometry part of the eigenproblem. The calculation bottleneck is the summary of mutual inductances from the reconstructed equilibrium current density. The with-mass spectrum can be made fully real-valued by the addition of a critical amount of damping with negligible effect on the vertical growth rate. Furthermore, the calculation has been implemented in the plasma control system at the DIII-D tokamak and used in experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Lattice quantum chromodynamics at large isospin density

We present an algorithm to compute correlation functions for systems with the quantum numbers of many identical mesons from lattice quantum chromodynamics (QCD). The algorithm is numerically stable and allows for the computation of n-pion correlation functions for n ϵ {1, … , N} using a single N × N matrix decomposition, improving on previous algorithms. We apply the algorithm to calculations of correlation functions with up to 6144 charged pions using two ensembles of gauge field configurations generated with quark masses corresponding to a pion mass m π = 170 MeV and spacetime volumes of (4.4 3 × 8.8) fm 4 and (5.8 3 × 11.6) fm 4 . We also discuss statistical techniques for the analysis of such systems, in which the correlation functions vary over many orders of magnitude. In particular, we observe that the many-pion correlation functions are well-approximated by log-normal distributions, allowing the extraction of the energies of these systems. Using these energies, the large-isospin-density, zero-baryon-density region of the QCD phase diagram is explored. A peak is observed in the energy density at an isospin chemical potential μ I ~ 1.5m π , signaling the transition into a Bose-Einstein condensed phase. The isentropic speed of sound, c s , in the medium is seen to exceed the ideal-gas (conformal) limit ($c^{2}_{s} ≤ 1/3)$ over a wide range of chemical potential before falling towards the asymptotic expectation at μ I ~ 15m π . These, and other thermodynamic observables, indicate that the isospin chemical potential must be large for the system to be well described by an ideal gas or perturbative QCD.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Numerical Circuit Synthesis and Compilation for Multi-State Preparation

Near-term quantum computers have significant error rates and short coherence times, so compilation of circuits to be as short as possible is essential. Two types of compilation problems are typically considered: circuits to prepare a given state from a fixed input state, called 'state preparation'; and circuits to implement a given unitary operation, for example by 'unitary synthesis'. In this paper we solve a more general problem: the transformation of a set of m states to another set of m states, which we call 'multi-state preparation'. State preparation and unitary synthesis are special cases; for state preparation, m=1, while for unitary synthesis, m is the dimension of the full Hilbert space. We generate and optimize circuits for multi-state preparation numerically. In cases where a top-down approach based on matrix decompositions is also possible, our method finds circuits with substantially (up to 40 %) fewer two-qubit gates. We discuss possible applications, including efficient preparation of macroscopic superposition ('cat') states and synthesis of quantum channels.

Szasz, Aaron↗

Performance of Compact Pulsed Thermal Imaging System for In-Service Applications. Pulsed thermal tomography nondestructive examination of additively manufactured reactor materials and components

Additive manufacturing (AM) is an emerging method for cost-efficient fabrication of complex topology nuclear reactor parts from high-strength corrosion resistance alloys, such as stainless steel and Inconel. AM of metallic structures for nuclear energy applications is currently based on laser powder bed fusion (LPBF) process, which has the capability of melting metallic powder and net shaping the structures with relatively high precision. Some of the challenges with using LPBF method for nuclear manufacturing include the possibility of introducing pores into metallic structures. Integrity of AM structures needs to be evaluated nondestructively because material flaws could lead to premature failures due to creep in high temperature nuclear reactor environment. Currently, there exist limited capabilities to evaluate actual AM structures nondestructively. Pulsed Thermography (PT) imaging provides a capability for non-destructive evaluation (NDE) of sub-surface defects in arbitrary size structures. The PT method is based on recording material surface temperature transients with infrared (IR) camera following thermal pulse delivered on material surface with flash light. The PT method has advantages for NDE of actual AM structures because the method involves one-sided non-contact measurements and fast processing of large sample areas captured in one image. The data cube of PT measurements consists of surface temperature taken at sequential time intervals T(x,y,t). Material defects can be detected either by analyzing the thermograms T(x,y,t) data cube, or by using thermal tomography (TT) algorithm to obtain 3D spatial reconstruction of thermal effusivity e(x,y,z). To reduce the cost and enable in-service NDE in spatially constrained environment, it is highly desirable to develop PT with compact and inexpensive IR camera. Following initial qualification of an AM component for deployment in a nuclear reactor, a compact PT system can also be used for in-service nondestructive evaluation (NDE) applications. However, data cube obtained with PT based on compact IR camera suffers from strong thermal noises and loss of features due to relatively low sampling rate. In this report we describe two unsupervised machine learning (ML) algorithms for enhancement of PT images obtained with compact IR camera. In one approach, we introduce Sparse Coding Discrete Cosine Transform (SC/DCT) algorithm to remove additive white Gaussian noise (AWGN) from spatial thermal effusivity reconstructions. In another approach we introduce a Spatial Temporal Denoised Thermal Source Separation (STDTSS) ML algorithm to process thermograms. The STDTSS algorithm consists of spatial and temporal denoising using Gaussian and Savitzky–Golay filtering, followed by the matrix decomposition using Principal Component Analysis (PCA), and Independent Component Analysis (ICA) to automatically detect flaws. In the work described in this report, we constructed a compact PT system using a relatively small and low-cost FLIR A65 camera, consisting on uncooled microbolometer detector. Performance of SC/DCT algorithm was demonstrated on enhancing TT images of Inconel 718 AM plate. Performance of the STDTSS methods was investigated using thermography data obtained from imaging stainless steel 316L specimens produced with LPBF method with imprinted calibrated porosity defects.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗