Search NASASearch

SEARCH · Search NASA

Results for “bit matrix”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A Study of Performance Portability of Low-bit Fused Matrix-Vector Multiplication Kernels in SYCL

Understanding the causes of performance gaps between a portable programming model and a vendor-specific programming model is important for improving performance portability. This paper studies performance portability of low-bit fused general matrix-vector multiplication kernels in SYCL on vendors’ graphics processing units (GPUs). This work introduces the use case, explains the kernel implementations in detail, evaluates the performance of the CUDA, HIP, and SYCL kernels on datacenter, desktop, and laptop GPUs, and investigates the causes of performance gaps. The results show that loop unrolling, kernel dispatch overhead, and sum reduction contribute to the gaps.

Jin, Zheming [ORNL] (ORCID:000000027197780X)

Coherence-Induced Deep Thermalization Transition in Random Permutation Quantum Dynamics

We report a phase transition in the projected ensemble—the collection of postmeasurement wave functions of a local subsystem obtained by measuring its complement. The transition emerges in systems undergoing random permutation dynamics, a type of quantum time evolution wherein computational basis states are shuffled without creating superpositions. It separates a phase exhibiting deep thermalization, where the projected ensemble is distributed over Hilbert space in a maximally entropic fashion (Haar random), from a phase where it is minimally entropic (“classical bit-string ensemble”). Crucially, this deep thermalization transition is invisible to the subsystem’s density matrix, which always exhibits thermalization to infinite temperature across the phase diagram. Through a combination of analytical arguments and numerical simulations, we show that the transition is tuned by the total amount of injected by the input state and the measurement basis, and is exhibited robustly across different microscopic models. Our findings represent a novel form of ergodicity-breaking universality in quantum many-body dynamics, characterized not by a failure of regular thermalization, but rather by a failure of deep thermalization.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Finding ways to reduce nuclear waste: searching for the unknown one step at a time

In my home country, Venezuela, research has been stagnant. Due to the political turmoil and the crisis, many educated people have left the country in search for a better life. This has caused a deficit in any technological and scientific advances, making Venezuela one of the first South American countries to have its rate of publications decline by 29% in 2013. Currently, Venezuela lacks the infrastructure and the means to keep up with the research progress as compared to other countries in South America, such as Brazil. Since coming to the United States (US), and currently working for a national laboratory, the active research environment endorses a wide range of careers and engineering programs that allow researchers to thrive at any given field. Researchers have access to funds and tools to succeed in developing materials for the future. There are 17 national laboratories in the US, and all of these have a different research focus/objective. As examples, Los Alamos National Laboratory and Sandia National Laboratory focus is on national homeland security, weapon science, radiation effects, among others. Argonne National Laboratory focuses on nuclear energy, energy storage, high performance computing, etc. At Idaho National Laboratory (INL) the research focuses on innovating nuclear energy and clean energy resources, critical infrastructure materials, along with fuel cycle solutions to manage, dispose and find ways to recycle current and future radiological waste. Compared to other national laboratories, INL focuses slightly more on applied processes and how nuclear energy can be innovated to next reactor design and technologies. The research being conducted at INL made me apply for a Seaborg distinguished postdoctoral position. For the position itself, the researcher must submit a proposal related to actinide chemistry on a research field area. In this position, 50% of my time will be focused on my own proposal. The proposal that I am working on is focused on the innovation of nuclear energy and fuel cycle recycling, which is why I was mainly interested on working at this national laboratory. To give a bit more context of what my proposal is about, a little bit of background is necessary: After the nuclear fuel (UO2) is used in a reactor, the fuel matrix is then characterized by various fission products (FP). Among these FP (including rare earth elements, alkali/alkaline earths, and actinides), many can potentially be recovered through nuclear reprocessing technologies. In pyroprocessing, the used nuclear fuel undergoes electrochemical dissolution into a molten chloride salt mixture in an electrorefiner. Initially, uranium is reduced onto an inert cathode by applied potentials. However, numerous remaining FPs accumulate in the melt and pose challenges for recovery by an inert electrode, particularly the rare earth elements (e.g., Nd, Gd, Pr, Sm) due to their multivalent oxidation states and tendencies toward side reactions, leading to their dissolution in the electrolyte. These recovery challenges result in inefficiencies and necessitate the continual discarding of the molten chloride salt, thereby generating additional waste. Furthermore, the presence of rare earth elements and other fission products in the molten salt electrolyte alters its physical and chemical properties, affecting both uranium recovery efficiency and the longevity of the molten chloride salt. To improve the recovery efficiency of the FP, specifically rare earth elements, I am investigating the fundamental interactions between rare earth elements in the molten chloride salt and their metallic form. The kinetic pathways and the chemical reactions of these elements will give insights on how the recovery efficiency can be improved. The interactions and speciation of these elements are being studied by spectro-electrochemistry at high temperature environments in quartz and other ceramic materials (e.g., alumina crucibles). Some of the challenges I am facing specifically relates the reactivity of some of these elements with different glass and crucible materials. Although my research focuses on fundamental science, it will benefit the applied process by generating new scientific knowledge and closing the gap for an efficient recycling of the waste: one step at a time.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

FTTN is a test suite to evaluate the numerical behaviors of matrix accelerators of GPUs (NVIDIA Tensor Cores and AMD Matrix Cores) in a quick and simple setting. Matrix accelerators are heavily used in today's computationally intense applications to speed up matrix multiplications. This test suite provides a comprehensive study on the numerical behaviors of these accelerators, including support for subnormals, rounding modes, extra precision bits and FMA features. Is there

Laguna Peralta, Ignacio

Robust Implicit Adaptive Low Rank Time-Stepping Methods for Matrix Differential Equations

In this work, we develop implicit rank-adaptive schemes for time-dependent matrix differential equations. The dynamic low rank approximation (DLRA) is a well-known technique to capture the dynamic low rank structure based on Dirac–Frenkel time-dependent variational principle. In recent years, it has attracted a lot of attention due to its wide applicability. Our schemes are inspired by the three-step procedure used in the rank adaptive version of the unconventional robust integrator (the so called BUG integrator) (Ceruti et al. in BIT Numer Math 62(4):1149–1174, 2022) for DLRA. First, a prediction (basis update) step is made computing the approximate column and row spaces at the next time level. Second, a Galerkin evolution step is invoked using an implicit solves for the small core matrix. Finally, a truncation is made according to a prescribed error threshold. Since the DLRA is evolving the differential equation projected on to the tangent space of the low rank manifold, the error estimate of the BUG integrator contains the tangent projection (modeling) error which cannot be easily controlled by mesh refinement. This can cause convergence issue for equations with cross terms. To address this issue, we propose a simple modification, consisting of merging the row and column spaces from the explicit step truncation method together with the BUG spaces in the prediction step. In addition, we propose an adaptive strategy where the BUG spaces are only computed if the residual for the solution obtained from the prediction space by explicit step truncation method, is too large. Here, we prove stability and estimate the local truncation error of the schemes under assumptions. We benchmark the schemes in several tests, such as anisotropic diffusion, solid body rotation and the combination of the two, to show robust convergence properties.

97 MATHEMATICS AND COMPUTING

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

While NVIDIA has been the dominant provider of GPUs for HPC and ML, now AMD has several offerings of GPUs. This encourages programmers to try out AMD GPUs for new codes and also port existing codes over. Unfortunately, without understanding the floating-point differences between these GPU types, software development or porting can introduce bugs—and currently such an understanding is lacking. The magnitude of this open question becomes clear if one imagines the the number of floating-point precision choices (FP16, FP32, etc.), floating-point formats (standard floats, brain-float, etc.), and execution units available (elementary units, matrix/tensor cores, etc.) Questions such as rounding modes and subnormal support are also important. Most of these answers are unknown today or are hard to access. We provide the first testing-guided approach that answers a significant number of these questions. We also devise tests to reveal internal information (e.g., extra bits kept) to make sure that our findings are reliable. Many of our tests employ systematically generated random-programs, others apply fast-math flags and some involve fused multiplyadd. Especially for tensor/matrix cores, the tests have nontrivial logic that we present Our testing approach is reusable for the plethora of GPUs yet to be introduced. Our findings include up to 7 ulps of difference between NVIDIA and AMD for sin and cos at FP32 precision and 3 ulp at FP64. In our study of matrix cores (NVIDIA) and tensor cores (AMD), we have extensively characterized rounding modes (truncation versus round-to-nearest), the number of extra internal bits kept (whether 3 bits are kept or not), subnormal support for inputs and outputs across four different floating-point formats and across NVIDIA A100 and AMD MI250X GPUs. We believe that this wealth of data becoming available for the first time may help avoid significant porting bugs when migrating code across these platforms.

Li, Xinyi

A 28 nm multiply-accumulate ASIC architecture for on-chip data compression in MHz frame rate X-ray and electron pixel detectors

Modern X-ray detector systems urgently require compact, efficient, and fast data compression schemes to handle the transmission of big data from pixel arrays, enabling frame rates in the MHz regime. Here, in this work, a data compression ASIC that implements a streaming fixed-length lossy compression scheme is introduced and analyzed, proving the feasibility and benefits of on-chip compression. The compression scheme utilizes a vector matrix product logic, which performs a number of floating-point multiplications, additions, and accumulations. The logic is verified, synthesized, and shown to fit in the area resource available for the X-ray detector under study, which comprises 192 × 168 pixels each of 12-bit width, and having a total area of 20 mm× 20 mm, about 2 mm× 20 mm of which are available for the digital logic. Several system architectures, precisions, and compression ratios ranging from 100 to 250 were analyzed to pave the way for on-chip fixed-length compression (e.g., principal component analysis, singular value decomposition) and data reduction (e.g., azimuthal integration) for X-ray and electron detectors.

Data compression

Inducing a tunable skyrmion-antiskyrmion system through ion beam modification of FeGe films

Abstract Skyrmions and antiskyrmions are nanoscale swirling textures of magnetic moments formed by chiral interactions between atomic spins in magnetic noncentrosymmetric materials and multilayer films with broken inversion symmetry. These quasiparticles are of interest for use as information carriers in next-generation, low-energy spintronic applications. To develop skyrmion-based memory and logic, we must understand skyrmion-defect interactions with two main goals—determining how skyrmions navigate intrinsic material defects and determining how to engineer disorder for optimal device operation. Here, we introduce a tunable means of creating a skyrmion-antiskyrmion system by engineering the disorder landscape in FeGe using ion irradiation. Specifically, we irradiate epitaxial B20-phase FeGe films with 2.8 MeV Au 4+ ions at varying fluences, inducing amorphous regions within the crystalline matrix. Using low-temperature electrical transport and magnetization measurements, we observe a strong topological Hall effect with a double-peak feature that serves as a signature of skyrmions and antiskyrmions. These results are a step towards the development of information storage devices that use skyrmions and antiskyrmions as storage bits, and our system may serve as a testbed for theoretically predicted phenomena in skyrmion-antiskyrmion crystals.

74 ATOMIC AND MOLECULAR PHYSICS

Quasiprobabilistic Readout Correction of Midcircuit Measurements for Adaptive Feedback via Measurement Randomized Compiling

Quantum measurements are a fundamental component of quantum computing. However, on present-day quantum computers, measurements can be more error prone than quantum gates and are susceptible to nonunital errors as well as nonlocal correlations due to measurement crosstalk. While readout errors can be mitigated in postprocessing, this is inefficient in the number of qubits due to a combinatorially large number of possible states that need to be characterized. In this work, we show that measurement errors can be tailored into a simple stochastic error model using randomized compiling, enabling the efficient mitigation of readout errors via quasiprobability distributions reconstructed from the measurement of a single preparation state in an exponentially large confusion matrix. We demonstrate the scalability and power of this approach by correcting readout errors without matrix inversion on a large number of different preparation states applied to a register of eight superconducting transmon qubits. Moreover, we show that this method can be extended to midcircuit measurements used for active feedback via quasiprobabilistic error cancellation, and we demonstrate the correction of measurement errors on an ancilla qubit used to detect and actively correct bit-flip errors on an entangled memory qubit. Our approach enables the correction of readout errors on large numbers of qubits and offers a strategy for correcting readout errors in adaptive circuits in which the results of midcircuit measurements are used to perform conditional operations on nonlocal qubits in real time.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Spectroscopic Signatures of Phonon Character in Molecular Electron Spin Relaxation

Spin–lattice relaxation constitutes a key challenge for the development of quantum technologies, as it destroys superpositions in molecular quantum bits (qubits) and magnetic memory in single molecule magnets (SMMs). Gaining mechanistic insight into the spin relaxation process has proven challenging owing to a lack of spectroscopic observables and contradictions among theoretical models. Here, we use pulse electron paramagnetic resonance (EPR) to profile changes in spin relaxation rates (T 1 ) as a function of both temperature and magnetic field orientation, forming a two-dimensional data matrix. For randomly oriented powder samples, spin relaxation anisotropy changes dramatically with temperature, delineating multiple regimes of relaxation processes for each Cu(II) molecule studied. We show that traditional T 1 fitting approaches cannot reliably extract this information. Single-crystal T 1 anisotropy experiments reveal a surprising change in spin relaxation symmetry between these two regimes. We interpret this switch through the concept of a spin relaxation tensor, enabling discrimination between delocalized lattice phonons and localized molecular vibrations in the two relaxation regimes. Variable-temperature T 1 anisotropy thus provides a unique spectroscopic method to interrogate the character of nuclear motions causing spin relaxation and the loss of quantum information.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Multistate resistance in TaN/(Hf,Zr)O 2 /Ta ferroelectric tunnel junctions

Ferroelectric tunnel junctions (FTJs) utilizing hafnium zirconium oxide (HZO) have emerged as promising non-volatile memory elements for microelectronics, compatible with back end of line (BEOL) complementary–metal–oxide semiconductor fabrication. This study investigates asymmetric electrode TaN/HZO/Ta devices with a 6 nm thick HZO layer as FTJs for multistate resistive memory applications. The individual FTJs exhibit a resistance ratio exceeding 10× when utilized as a binary state device, with pulsing between −1.7 and +1.4 V to set the high resistance state (HRS) and low resistance state (LRS), respectively. Following with reduced write voltage pulses allows the ferroelectric device to operate with a selection of over 32 distinct resistance states (2 5 bits) between the LRS and HRS. This work then explores the stability of the resistance states during write/read pulse cycling, along with the stability of the state after multiple read pulses. Accessing the multibit state shows stability within 50 reads with the binary state remaining stable for more than 4000 reads pulses. With their multistate tunability and versatility, FTJs hold promise as BEOL memory elements for compute-in-memory (CiM) arrays, binary digital memory, or weighted vector matrix multiplication applications with low power consumption during computations.

CMOS