Search NASA⌕ Search

SEARCH · Search NASA

Results for “tensor program generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Cross-Feature Transfer Learning for Efficient Tensor Program Generation

Tuning tensor program generation involves navigating a vast search space to find optimal program transformations and measurements for a program on the target hardware. The complexity of this process is further amplified by the exponential combinations of transformations, especially in heterogeneous environments. This research addresses these challenges by introducing a novel approach that learns the joint neural network and hardware features space, facilitating knowledge transfer to new, unseen target hardware. A comprehensive analysis is conducted on the existing state-of-the-art dataset, TenSet, including a thorough examination of test split strategies and the proposal of methodologies for dataset pruning. Leveraging an attention-inspired technique, we tailor the tuning of tensor programs to embed both neural network and hardware-specific features. Notably, our approach substantially reduces the dataset size by up to 53% compared to the baseline without compromising Pairwise Comparison Accuracy (PCA). Furthermore, our proposed methodology demonstrates competitive or improved mean inference times with only 25–40% of the baseline tuning time across various networks and target hardware. The attention-based tuner can effectively utilize schedules learned from previous hardware program measurements to optimize tensor program tuning on previously unseen hardware, achieving a top-5 accuracy exceeding 90%. This research introduces a significant advancement in autotuning tensor program generation, addressing the complexities associated with heterogeneous environments and showcasing promising results regarding efficiency and accuracy.

97 MATHEMATICS AND COMPUTING↗

Interfacing electron and neutrino quasielastic scattering cross sections with the spectral function in GENIE

Progress in neutrino-nucleus cross section models is being driven by the need for highly accurate predictions for the neutrino oscillation community. These sophisticated models are being developed within a microscopic description of the nucleus with the goal of encompassing all reaction modes relevant for the accelerator neutrino program. The disconnect between these microscopic models and the event generators that will be used in the next generation of experiments represents a critical obstacle that must be overcome in order to precisely measure the neutrino oscillation parameters. To this end we have developed a hadron tensor interface for lepton-nucleus quasielastic (QE) scattering within the GENIE event generator as a proof of principle, with the broader goal of creating an efficient pipeline for incorporating advanced theoretical models in event generators. As a demonstration of this interface we have implemented the spectral function model into GENIE by connecting theorist provided fortran code through the hadron tensor interface. The spectral function model offers a more complete description of the nuclear ground state, as well as the ability to provide quantifiable theoretical uncertainties. Finally, we validate this implementation and compare its predictions against data and against QE models already available in GENIE.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

While NVIDIA has been the dominant provider of GPUs for HPC and ML, now AMD has several offerings of GPUs. This encourages programmers to try out AMD GPUs for new codes and also port existing codes over. Unfortunately, without understanding the floating-point differences between these GPU types, software development or porting can introduce bugs—and currently such an understanding is lacking. The magnitude of this open question becomes clear if one imagines the the number of floating-point precision choices (FP16, FP32, etc.), floating-point formats (standard floats, brain-float, etc.), and execution units available (elementary units, matrix/tensor cores, etc.) Questions such as rounding modes and subnormal support are also important. Most of these answers are unknown today or are hard to access. We provide the first testing-guided approach that answers a significant number of these questions. We also devise tests to reveal internal information (e.g., extra bits kept) to make sure that our findings are reliable. Many of our tests employ systematically generated random-programs, others apply fast-math flags and some involve fused multiplyadd. Especially for tensor/matrix cores, the tests have nontrivial logic that we present Our testing approach is reusable for the plethora of GPUs yet to be introduced. Our findings include up to 7 ulps of difference between NVIDIA and AMD for sin and cos at FP32 precision and 3 ulp at FP64. In our study of matrix cores (NVIDIA) and tensor cores (AMD), we have extensively characterized rounding modes (truncation versus round-to-nearest), the number of extra internal bits kept (whether 3 bits are kept or not), subnormal support for inputs and outputs across four different floating-point formats and across NVIDIA A100 and AMD MI250X GPUs. We believe that this wealth of data becoming available for the first time may help avoid significant porting bugs when migrating code across these platforms.

Li, Xinyi↗

Applying Quantum Computing to Simulate Power System Dynamics

Power system dynamics are generally modeled by high dimensional nonlinear differential-algebraic equations due to a large number of generators, loads, and transmission lines. Thus, its computational complexity grows exponentially with the system size. This paper demonstrates the potential use of quantum computing algorithms to model the power system dynamics. Leveraging a symbolic programming framework, we equivalently convert the power system dynamics’ differential algebraic equations (DAEs) into ordinary differential equations (ODEs), where the data of the state vector can be encoded into quantum computers via amplitude encoding. The system's nonlinearity is captured by Taylor polynomial expansion, the quantum state tensor, and Hamiltonian simulation, whereas state variables can be updated by a quantum linear equation solver. Our results show that quantum computing can simulate the dynamics of the power system with high accuracy, whereas its complexity is polynomial in the logarithm of the system dimension. Our work also illustrates the use of scientific machine learning tools for implementing scientific computing concepts, e.g., Taylor expansion, DAEs/ODEs transform, and quantum computing solver, in the field of power engineering.

Tran, Huynh↗

Geometric Interpretation of the Cluster Location Problem Part II: Application to the Pahala, Hawaii, Earthquake Sequence

In the companion “Theory” article, we presented a new framing of the seismic location problem in terms of differential geometry (Harris et al., 2025). From that viewpoint, we developed a “project and correct” approach for estimating the relative locations of earthquakes. Here, in this study, we use project and correct to estimate high-precision relative locations of events from an earthquake sequence beneath the town of Pahala, Hawaii, using high-precision correlation-derived picks. The sequence was active from 2020 through 2022 and produced many highly correlated signals at Hawaii Volcano Observatory (HVO) stations on the island of Hawaii. The data we inverted consisted of 2882 events with observations at 5 HVO stations. For comparison with the travel-time image, we also produced conventional hypocenter solutions using both the Bayesloc program (Myers et al., 2007, 2009) and a purpose-built double-difference code. There were obvious structural elements in the resulting image, the resolution of which we used to test the performance of the project and the correct algorithm. For the projection step, we first produced a 3D local basis using an singular value decomposition (SVD) of the 2882 groups of times. Projection of the travel-time vectors into this basis resulted in an image with structures similar to those produced by our conventional locators, but with distortion as predicted by theory. Removing the distortion requires an inverse operator generated from the metric tensor at the geometric centroid of the events. We compared two approaches to obtaining such an inverse operator. The first uses an estimate of the geographic centroid of the event cloud from the centroid of the travel-time data. The second approach uses the centroid of the conventionally produced locations. The first approach produces a corrected image very similar to the conventional results, but with a rotation. The corrected image produced using the conventionally derived centroid is a near-exact match to the conventional locations.

Dodge, Douglas A. [Lawrence Livermore National Lab↗

Design, Control and Application of Next Generation Qubits

Design, Control and Application of Next Generation Qubits Arun Bansil, Northeastern University (Principal Investigator) Claudio Chamon, Boston University (Co-Investigator) Adrian Feiguin, Northeastern University (Co-Investigator) Liang Fu, MIT (Co-Investigator) Eduardo Mucciolo, Univ. of Central Florida (Co-Investigator) Qimin Yan, Temple University (Co-Investigator) The quest for developing technologies for manipulating and storing information quantum mechanically is currently led by approaches that include Josephson-junctions, ion-traps, and qubits generated by defect spins in solids. Topological qubits, however, are inherently more robust to decoherence by environmental effects, and should be able to sprint ahead once practical barriers have been overcome. At the present stage of the development of the field, it is important to explore a variety of architectures and materials beyond the conventional paradigms in order to seed breakthroughs toward building a scalable quantum computer. Our comprehensive theoretical research program involved four interconnected thrusts as follows. • A materials discovery effort in two-dimensional compounds in search of materials to support Majorana zero modes and defect structures suitable as qubits. • Exploration of architectures for topological quantum computation by investigating both superconducting Majorana qubits, and robust platforms for braiding with new “meta-materials” built of arrays of Majorana qubits. • Investigation of properties of hybrid metal-organic qubits based on transition-metal centers in graphene, and molecular crystals of polyaromatic complexes with embedded transition-metal atoms. • Development of tensor-network and semiclassical approaches to study decoherence in the presence of random and dispersive spin baths, and NV centers in diamond. The full spectrum of theoretical and numerical approaches was used to address the goals of this project including first-principles, density-matrix-renormalization group, tensor networks, and data-driven high-throughput approaches using materials database and machine-learning.

36 MATERIALS SCIENCE↗