Search NASA⌕ Search

SEARCH · Search NASA

Results for “Arithmetic”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Reconfigurable neuromorphic components and algorithms for next-generation artificial intelligence

Digital transistor-based general-purpose hardware (e.g., central processing units) is the dominant solution to support both traditional computing (logic, arithmetic, etc.) as well as modern artificial intelligence. State-of-the-art research has shown feasibility of post-digital physics-based neuromorphic hardware, which is hypothesized to support artificial intelligence algorithms with orders-of-magnitude improved time/energy efficiencies. But such research has not been widely deployed mainly because of such novel hardware’s extreme application-specificity, and the dominance of low-cost general-purpose (but inefficient) digital hardware. To make use of the novel algorithms and the superlative performance of physics-based hardware, we need to identify scientific principles that can enable generality in physics-based hardware. This work resulted in two important broad outcomes – first, we demonstrate fully reconfigurable neuromorphic components, and second, we demonstrate a viable artificial intelligence learning algorithm that can exploit the functioning of neuromorphic hardware. We demonstrate up to five orders of magnitude improvement in energy efficiency compared to the best general-purpose digital hardware.

97 MATHEMATICS AND COMPUTING↗

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING↗

Controlling Oxidation of Nb in Oxygen Abundant Environments

Modern particle accelerators depend on Superconducting Radio Frequency (SRF) cavities made from high-purity niobium (Nb) to achieve optimal performance, including high quality factors and strong accelerating gradients. However, when exposed to air, niobium naturally forms a complex oxide layer that can introduce surface imperfections and carbon contamination. This work examines an alternative oxidation strategy under oxygen-rich conditions to better regulate the oxide formation process. The ultimate objective is to improve surface uniformity and cleanliness, thereby reducing defect density and enhancing performance. We used Confocal Microscopy, Scanning Electron Microscopy (SEM), and X-Ray Photoelectron Spectroscopy (XPS) to analyze surface changes. We used standard metrics, like Arithmetic Average Roughness (Ra) and Root Mean Square (Rq). Three oxidation methods were applied: short, extended, and a repeated HF and H₂O₂ oxidation process. Preliminary results show promising reductions in carbon contamination and surface defects.

Romero, Juan [Fermilab]↗

Infinite quantum signal processing

Quantum signal processing (QSP) represents a real scalar polynomial of degree d using a product of unitary matrices of size 2 × 2 , parameterized by ( d + 1 ) real numbers called the phase factors. This innovative representation of polynomials has a wide range of applications in quantum computation. When the polynomial of interest is obtained by truncating an infinite polynomial series, a natural question is whether the phase factors have a well defined limit as the degree d → ∞ . While the phase factors are generally not unique, we find that there exists a consistent choice of parameterization so that the limit is well defined in the ℓ 1 space. This generalization of QSP, called the infinite quantum signal processing, can be used to represent a large class of non-polynomial functions. Our analysis reveals a surprising connection between the regularity of the target function and the decay properties of the phase factors. Our analysis also inspires a very simple and efficient algorithm to approximately compute the phase factors in the ℓ 1 space. The algorithm uses only double precision arithmetic operations, and provably converges when the ℓ 1 norm of the Chebyshev coefficients of the target function is upper bounded by a constant that is independent of d . This is also the first numerically stable algorithm for finding phase factors with provable performance guarantees in the limit d → ∞ .

Dong, Yulong [Department of Mathematics, Universit↗

Computer program provides linear sampled- data analysis for high order systems

Computer program performs transformations in the order S-to W-to Z to allow arithmetic to be completed in the W-plane. The method is based on a direct transformation from the S-plane to the W-plane. The W-plane poles and zeros are transformed into Z-plane poles and zeros using the bilinear transformation algorithm.

Bunn, D. B.↗

Digital data averager improves conventional measurement system performance

Multipurpose digital averager provides measurement improvement in noisy signal environments. It provides increased measurement accuracy and resolution to basic instrumentation devices by an arithmetical process in real time. It is used with standard conventional measurement equipment and digital data printers.

Naylor, T. K.↗

Bounds for the Horner sums.

Chebyshev polynomials maximum property, examining effect of roundoff errors in Horner scheme for floating point arithmetic

Reimer, M.↗

Self testing and repairing computer - A concept

STAR computer has five redundant modular function units, fixed store, arithmetic, memory, input, and output. Each unit is connected to a diagnostic control unit, each is coded for error detection and error correction. Separation into function units permits assembly of many different systems from the set of units.

Avizienis, A. A.↗

Computer

Arithmetic and code-checking routines of ILLAR system and recursive subprogram in FORTRAN compiler

Bouknight, J.↗

Efficient digital comparison technique for logic circuits

Tolerance compare technique indicates discompare only when numerical difference value exceeds prescribed limit. Algorithm involving binary number properties is defined, in lieu of arithmetic operation which requires relatively complex circuitry. Extension of algorithm may be made to encompass tolerances other than one unit.

Mccarthy, C. E.↗

Delta modulation

The conclusions of the design research of the song adaptive delta modulator are presented for source encoding voice signals. The variation of output SNR vs input signal power/when 8, 9, and 10 bit internal arithmetic is employed. Voice intelligibility tapes to test the 10-bit system are used. An analysis of a delta modulator is also presented designed to minimize the in-band rms error. This is accomplished by frequency shaping the error signal in the modulator prior to hard limiting. The result is a significant increase in the output SNR measured after low pass filtering.

Schilling, D. L.↗