Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer Hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Development of a Multi-Robot System for Autonomous Inspection of Nuclear Waste Tank Pits

This paper introduces the overall design plan, development timeline, and preliminary progress of the Autonomous Pit Exploration System project. This project aims to develop an advanced multi-robot system for the efficient inspection of nuclear waste-storage tank pits. The project is structured into three phases: Phase 1 involves data collection and interface definition in collaboration with Hanford Site experts and university partners, focusing on tank riser geometry and hardware solutions. Phase 2 includes the selection of sensors and robot components, detailed mechanical design, and prototyping. Phase 3 integrates all components into a cohesive system managed by a master control package which also incorporates digital twin and surrogate models, and culminates in comprehensive testing and validation at a simulated tank pit at the Idaho National Laboratory. Additionally, the system’s communication design ensures coordinated operation through shared data, power, and control signals. For transportation and deployment, an electric vehicle (EV) is chosen to support the system for a full 10 h shift with better regulatory compliance for field deployment. A telescopic arm design is selected for its simple configuration and superior reach capability and controllability. Preliminary testing utilizes an educational robot to demonstrate the feasibility of splitting computational tasks between edge and cloud computers. Successful simultaneous localization and mapping (SLAM) tasks validate our distributed computing approach. More design considerations are also discussed, including radiation hardness assurance, SLAM performance, software transferability, and digital twinning strategies.

Nuclear waste management↗

Addressing Rising Energy Demand Through Innovation

The U.S. is facing a significant increase in energy demand, driven by AI advancements, the rapid expansion of data centers, manufacturing and industrial growth, and the electrification of transportation and buildings. Buildings alone account for approximately 75% of U.S. electricity consumption and 40% of total energy use. To address these challenges, NLR leverages its state-of-the-art research facilities, advanced energy modeling, hardware-in-the-loop emulation, and real-world demonstrations to provide data-driven insights that de-risk emerging energy solutions, increase efficiency and demand flexibility, optimize grid controls, and identify vulnerabilities to enhance energy security. This presentation will highlight our research ecosystem and its role in supporting a more reliable, affordable, and adaptive energy infrastructure in the face of accelerating demand.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Solving reaction dynamics with quantum computing algorithms

The description of quantum many-body dynamics is extremely challenging on classical computers, as it can involve many degrees of freedom. However, the time evolution of quantum states is a natural application for quantum computers that are designed to efficiently perform unitary transformations. Here, in this paper, we study quantum algorithms for response functions, relevant for describing different reactions governed by linear response. We focus on nuclear-physics applications and consider a qubit-efficient mapping on the lattice, which can efficiently represent the large volumes required for realistic scattering simulations. For the case of a contact interaction, we develop an algorithm for time evolution based on the Trotter approximation that scales logarithmically with the lattice size and is combined with quantum phase estimation. We eventually focus on the nuclear two-body system and a typical response function relevant for electron scattering as an example. We also investigate ground-state preparation and examine the total circuit depth required for a realistic calculation and the hardware noise level required to interpret the signal.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

TChem-atm (v2.0.0): scalable performance-portable multiphase atmospheric chemistry

We present TChem-atm, a performance-portable approach that enables efficient simulation of chemically detailed and multiphase atmospheric chemistry on modern heterogeneous computing architectures. Unlike previous efforts that rely on architecture-specific code or focus exclusively on gas-phase chemistry, TChem-atm supports fully coupled gas–aerosol systems with execution across CPUs, NVIDIA GPUs, and AMD GPUs through the Kokkos programming model. It integrates the flexible multiphase capabilities of the Community Atmospheric Model Chemistry Package (CAMP) with the high-performance kinetic routines of TChem, and includes automatic Jacobian construction with support for a range of stiff ODE solvers. In a proof-of-concept integration with the particle-resolved model PartMC, TChem-atm reproduces the existing PartMC–CAMP implementation within solver tolerances and delivers substantial GPU speedups, especially for large particle populations. Performance benchmarks reveal substantial speedups on GPU platforms, particularly for large particle populations, with consistent results across hardware backends. TChem-atm enables performance-portable execution across CPUs and GPUs, though optimal efficiency may require modest architecture-specific tuning (e.g., team and vector sizes), with up to a twofold improvement on the NVIDIA H100. It directly supports sectional and particle-resolved host models, while modal aerosol schemes require minor adaptation to provide particle-scale quantities such as representative diameters. By enabling chemically detailed, multiphase simulations with performance portability and host-model flexibility, TChem-atm facilitates the incorporation of advanced chemistry into atmospheric models.

Díaz-Ibarra, Oscar Homero [Sandia National Laborat↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗

A Suppression-based STDP Rule Resilient to Jitter Noise in Spike Patterns for Neuromorphic Computing

Multi-spike models of synaptic plasticity, such as the triplet and suppression spike-timing-dependent plasticity (STDP) rules, exhibit better alignment with neurophysiological data in the brain compared to the pair-based STDP rule. Previous studies have empirically shown that the pair-based STDP rule can detect spatiotemporal spike patterns hidden in equally dense distractor spike trains in an unsupervised manner. However, it fails to detect spike patterns influenced by jitter noise. Given that spiking neural networks (SNNs) exhibit variability in generated spike trains in response to the same inputs, it becomes imperative to have learning rules capable of detecting spike patterns even in the presence of jitter noise. In this study, we introduce a simplified suppression-based STDP rule that demonstrates significantly enhanced tolerance to jitter in spike patterns compared to the pair-based STDP rule. Unlike the ideal suppression STDP rule, characterized by an exponential learning window and requiring high-resolution synapses, the simplified rule limits the synaptic efficacy update to a single bit at any given instant. Moreover, it employs 4-bit fixed-point synapses, facilitating straightforward implementation in neuromorphic hardware.

Gautam, Ashish [ORNL]↗

Sparse non-Markovian Noise Modeling of Transmon-Based Multi-Qubit Operations

The influence of noise on quantum dynamics is one of the main factors preventing current quantum processors from performing accurate quantum computations. Sufficient noise characterization and modeling can provide key insights into the effect of noise on quantum algorithms and inform the design of targeted error protection protocols. However, constructing effective noise models that are sparse in model parameters, yet predictive can be challenging. In this work, we present an approach for effective noise modeling of multi-qubit operations on transmon-based devices. Through a comprehensive characterization of seven devices offered by the IBM Quantum Platform, we show that the model can capture and predict a wide range of single- and two-qubit behaviors, including non-Markovian effects resulting from spatiotemporally correlated noise sources. The model’s predictive power is further highlighted through multi-qubit dynamical decoupling demonstrations and an implementation of the variational quantum eigensolver. As a training proxy for the hardware, we show that the model can predict expectation values within a relative error of 0.5%; this is a sevenfold improvement over default hardware noise models. Through these demonstrations, we highlight key error sources in superconducting qubits and illustrate the utility of reduced noise models for predicting hardware dynamics.

open quantum systems & decoherence↗

Use of Legacy Maritime Protocols Increases Exploitability of Virtual Aids to Navigation

With increased reliance on Virtual Aid(s) to Navigation (VAtoN) - also known as electronic Aid(s) to Navigation (eAtoN), or virtual buoys - a cyber event is likely to cause disruption to international maritime shipping. VAtoN has no physical hardware for visual reference and displays only on a vessel’s Electronic Chart Display Information System (ECDIS) and Automatic Radar Plotting Aid (ARPA); therefore, mariners must rely on the accuracy of the information provided. As VAtoN uses the National Maritime Electronics Association (NMEA) 0183 protocol for both Global Navigation Satellite System (GNSS) and Automatic Identification System (AIS), an insecure protocol that has been proven susceptible to spoofing, denial, and manipulation, the likelihood of a cyber-related event increases substantially.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Reactor System Demonstration with Cyber-Attack Scenarios Using CrowPis and Arduino Microcontrollers

This study covers developing and simulating nuclear reactor system using CrowPis and Arduino microcontrollers for demonstrating cyber-attack scenarios. The team was tasked with implementing more sensors and cybersecurity aspects to the reactor program that was created by last year’s high school interns. The team received the opportunity to collaborate and obtain advice from multiple university interns that helped us gain a better perspective of our project. Our mentor’s background in nuclear science was pivotal to our understanding what we could add to the reactor program to make it as realistic as possible. The first week of our internship was spent reading as much material as possible to gain an understanding of and the background for cyber-attacks and nuclear science. Nuclear science was a new horizon for each of the high school interns on the team, so spending this time in the beginning of our internship was crucial to our success. For the remaining portion of our internship, the team collectively did our best to implement as many sensors and use as much hardware as we could to make an accurate representation of a nuclear reactor in the program that was created. This internship was a big learning experience for everyone on the team. We all gained so many insights into the nuclear world and how it can benefit our lives, as well as how so many moving pieces are needed for it to work properly.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗

SineKAN: Kolmogorov-Arnold Networks using sinusoidal activation functions

Recent work has established an alternative to traditional multi-layer perceptron neural networks in the form of Kolmogorov-Arnold Networks (KAN). The general KAN framework uses learnable activation functions on the edges of the computational graph followed by summation on nodes. The learnable edge activation functions in the original implementation are basis spline functions (B-Spline). Here, we present a model in which learnable grids of B-Spline activation functions are replaced by grids of re-weighted sine functions (SineKAN). We evaluate numerical performance of our model on a benchmark vision task. We show that our model can perform better than or comparable to B-Spline KAN models and an alternative KAN implementation based on periodic cosine and sine functions representing a Fourier Series. Further, we show that SineKAN has numerical accuracy that could scale comparably to dense neural networks (DNNs). Compared to the two baseline KAN models, SineKAN achieves a substantial speed increase at all hidden layer sizes, batch sizes, and depths. Current advantage of DNNs due to hardware and software optimizations are discussed along with theoretical scaling. Additionally, properties of SineKAN compared to other KAN implementations and current limitations are also discussed.

Reinhardt, Eric↗

An electro-optical Mott neuron based on niobium dioxide

Various applications—including brain-like computing and on-chip artificial vision—increasingly demand a combination of electronic and photonic techniques. However, integrating both approaches on a single chip is challenging, and solutions typically rely on disparate components with power-hungry signal conversions. Here, in this paper, we report electro-optical Mott neurons that combine visible light emission with electrical threshold switching, as well as neuron-like oscillations. The devices are based on thin films of sputtered niobium dioxide (NbO 2 ), a Mott insulator–metal transition material, operating at room temperature and emitting light that peaks around 810 nm. Operando measurements reveal an electronic origin to the light emission: charge carrier relaxation initiated by high-field transport in the NbO 2 . Our devices combine electrical and optical functions within a single material, thereby expanding the options available for future artificial intelligence hardware.

electrical engineering↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Mitigating the Effects of Au-Al Intermetallic Compounds Due to High-Temperature Processing of Surface-Electrode Ion Traps

Stringent physical requirements need to be met for the high-performing surface-electrode ion traps used in quantum computing and timekeeping. In particular, these traps must survive a high-temperature environment for vacuum chamber preparation and support high RF voltage on closely spaced electrodes. Due to the use of gold wire bonds on aluminum pads, intermetallic growth can lead to wire bond failure via breakage or high resistance, limiting the lifetime of a trap assembly to a single multiday bake at 200 ° C. Using traditional thick metal stacks to prevent intermetallic growth, however, can result in trap failure due to RF breakdown events. Through high-temperature experiments, we conclude that an ideal metal stack for ion traps is Ti/Pt/Au (20/100/250 nm), which allows for a cumulative bakeable time of roughly 86 days without compromising the trap voltage performance. This increase in the bakeable lifetime of ion traps will remove the need to discard otherwise functional ion traps when vacuum hardware is upgraded, which will greatly benefit ion trap experiments.

Haltli, Raymond A.↗

Speeding Up Hartree–Fock in JuliaChem with Density Fitting

In this work, the density fitting (DF) approximation is added to the restricted Hartree–Fock (RHF) implementation in the JuliaChem computational chemistry code. Utilizing a DF algorithm that uses symmetry and integral screening, a significant reduction in time to compute the Fock matrix is achieved. The symmetry and screening DF-RHF techniques were adapted to be performed on graphics processing units (GPUs), which are well suited to perform the matrix multiplications that comprise the bulk of the Fock build time in DF-RHF. The JuliaChem DF-RHF GPU algorithm employs a novel approach that automatically switches between two DF-RHF algorithms depending on the number of basis functions in the calculation. The JuliaChem GPU DF-RHF implementation demonstrates up to 2× speedup for Fock build times compared to the existing best-in-class GPU DF-RHF implementation by operating directly on screened intermediate matrices. Due to the high portability of the Julia language code, the JuliaChem CPU and GPU DF-RHF implementations could be benchmarked on a variety of CPU and GPU architectures from multiple hardware vendors.

Hayes, John J. [Ames Laboratory, and Iowa State Un↗

Triangular cross-section beam splitters in silicon carbide for quantum information processing

Abstract Triangular cross-section color center photonics in silicon carbide is a leading candidate for scalable implementation of quantum hardware. Within this geometry, we model low-loss beam splitters for applications in key quantum optical operations such as entanglement and single-photon interferometry. We consider triangular cross-section single-mode waveguides for the design of a directional coupler. We optimize parameters for a 50:50 beam splitter. Finally, we test the experimental feasibility of the designs by fabricating triangular waveguides in an ion beam etching process and identify suitable designs for short-term implementation.

42 ENGINEERING↗

Oak Ridge National Laboratory Evaluation of Stream-Trained Models in Practice

The goal of this integration is to replicate the results from the original paper Autonomous Utility Pole Identification on different camera hardware and integrate the model into a live video stream provided by the unmanned aerial system (UAS) itself while in operation. This involves retraining the original model and validating its efficacy on multiple camera modules to select the most effective device for installation. Moreover, this integration requires writing software to handle the reception of a real-time streaming protocol stream from the UAS and run each frame through the model while allowing a user to monitor the camera feed.

97 MATHEMATICS AND COMPUTING↗

Evaluating the Limits of QAOA Parameter Transfer at High-Rounds on Sparse Ising Models With Geometrically Local Cubic Terms

The emergent practical applicability of the Quantum Approximate Optimization Algorithm (QAOA) for approximate combinatorial optimization is a subject of considerable interest. One of the primary limitations of QAOA is the task of finding a set of good parameters, which is usually done using a variational optimization loop. Parameter transfer, or parameter concentration, is a phenomenon where QAOA angles trained on problem instances that are self-similar tend to perform well for other problem instances from that similar class. This suggests a potentially highly efficient and scalable non-variational learning method for QAOA angle finding. In this work, we systematically study QAOA parameter transferability from small problem sizes (16 and 27 decision variables) onto large problem instances (up to 156 qubits) for heavy-hex graph Ising models with geometrically local higher order terms using the Julia based QAOA simulation tool \texttt{JuliQAOA} to perform classical angle finding for up to $49$ QAOA layers ($p$). Parameter transfer of the fixed angles is validated using a combination of full statevector, Projected Entangled Pair States (PEPS), Matrix Product State (MPS), and LOWESA numerical simulations. We find that the QAOA parameter transfer from single instances applied to other (unseen) problem instances does not in general provide monotonically improving performance as a function of $p$ - there are many cases where the performance temporarily decreases as a function of $p$ - but despite this the transferred angles have a general trend of improved expectation value as the QAOA depth increases, in many cases converging close to the true ground-state energy of the $100+$ qubit instances. We also sample the hardware-compatible Ising models using the ensemble of transfer-learned QAOA parameters on several superconducting qubit IBM Quantum processors with 127, 133, and 156 qubits. We find continuous solution quality improvement of the hardware-compatible QAOA circuits run on the IBM NISQ processors up to $p=5$ on \texttt{ibm\_fez}, up to $p=9$ on \texttt{ibm\_torino}, and up to $p=10$ on \texttt{ibm\_pittsburgh}.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗