Search NASA⌕ Search

SEARCH · Search NASA

Results for “computer hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Securing Grid-interactive Efficient Buildings (GEB) through Cyber Defense and Resilient System (CYDRES)

The DOE CYDRES project is driven by the urgent need to address critical research gaps in the domain of cyber-physical security of smart buildings, including Grid-interactive Efficient Buildings (GEBs). CYDRES, a real-time advanced building resilient platform, aims to enhance the cyber-attack-immune capabilities of buildings through multi-layered prevention, detection, and adaptation mechanisms. CYDRES consists of five key modules: a multi-layer network analyzer, an Automatic Fault Detection, Diagnosis, and Prognosis (AFDDP) framework, an intelligent mode selector, a cyber-resilient control framework, and a situation awareness platform. The Network Analyzer employs a data-driven framework that includes a protocol state learning tool and a CRF (Conditional Random Field) command validator. In Hardware-In-the-Loop (HIL) testbeds, it achieved 100% detection accuracy with a false alarm rate of 3%, validating its efficacy in identifying selected cyber-attacks. The AFDDP framework leverages pattern matching, PCA (Principal Component Analysis)-based strategies, and a DBN (Dynamic Bayesian Network)-based fault diagnosis approach to pinpoint the causes of physical system abnormalities using Building Automation System (BAS) data. In HIL experiments, the AFDDP module attained a detection accuracy of over 95% with a false alarm rate below 7%. Additionally, the fault detector utilized machine learning (Random Forest) and deep learning (Multi-Layer Perceptron) methods with acoustic sensor data to achieve a 100% fault detection accuracy in Heating, Ventilation, and Air-Conditioning (HVAC) equipment. The Mode Selector offered real-time impact analysis, allowing immediate actions to protect BASs in the face of emerging threats. The cyber-resilient control framework included an adaptive Model Predictive Control (MPC) and a measurement compensator, reducing temperature violations by up to 94% and improving the total demand flexibility by up to 70% in HIL experiments. Such HIL experiments covered a cyber-attack case and a physical fault case, showcasing CYDRES’ efficiency in maintaining operational continuity during threats. The situation awareness platform in Grafana enhanced real-time threat detection and response visualization, augmenting the operational awareness for building operators. CYDRES demonstrated high technical effectiveness in various test scenarios, particularly in HIL environments. The project's phased development approach ensured efficient use of resources, highlighting its practical feasibility and readiness for commercialization. By enhancing the security and resilience of building operations, CYDRES represents a significant advance in mitigating risks associated with cyber-physical systems, thereby enhancing public confidence in the safety of modern building infrastructure. Future directions for the project include expanding testing protocols, refining AFDDP methodologies, exploring more comprehensive resilient control strategies, and testing in real commercial buildings.

42 ENGINEERING↗

Synchronization for CXL Based Memory

Compute Express Link (CXL) is an important emerging standard for disaggregated memory. While this standard provisions coherency across numerous hosts and devices, implementing hardware support for type three devices is challenging. In this work, we look at the overhead of software synchronization and using software-based coherency. Moreover, we discuss the limits of software-based coherency in fully expressing modern synchronization techniques for a CXL-based disaggregate memory system. We demonstrate our approach using a CXL hardware prototype and running a version of the famous Peterson Lock (enhanced to run with more than two threads). We analyze its performance and share how more advanced synchronization techniques might interact with software-based coherence CXL hardware and program execution models.

High Performance Computing (HPC)↗

xesn: Echo state networks powered by Xarray and Dask

Xesn is a Python package that allows scientists to easily design Echo State Networks (ESNs) for forecasting problems. ESNs are a Recurrent Neural Network architecture introduced by Jaeger (2001) that are part of a class of techniques termed Reservoir Computing. One defining characteristic of these techniques is that all internal weights are determined by a handful of global, scalar parameters, thereby avoiding problems during backpropagation and reducing training time significantly. Because this architecture is conceptually simple, many scientists implement ESNs from scratch, leading to questions about computational performance. Xesn offers a straightforward, standard implementation of ESNs that operates efficiently on CPU and GPU hardware. The package leverages optimization tools to automate the parameter selection process, so that scientists can reduce the time finding a good architecture and focus on using ESNs for their domain application. Importantly, the package flexibly handles forecasting tasks for out-of-core, multi-dimensional datasets, eliminating the need to write parallel programming code. Xesn was initially developed to handle the problem of forecasting weather dynamics, and so it integrates naturally with Python packages that have become familiar to weather and climate scientists such as Xarray (Hoyer & Hamman, 2017). However, the software is ultimately general enough to be utilized in other domains where ESNs have been useful, such as in signal processing (Jaeger & Haas, 2004).

97 MATHEMATICS AND COMPUTING↗

Accelerating computing for the future electric grid (CRADA Final Report)

As a participant in the Cyclotron Road Lab-Embedded Entrepreneurship Program (LEEP), Vellex Computing, Inc. has successfully validated the "Vellex Computing Stack," a breakthrough Analog Neural Computer (ANC) specifically designed for high-performance edge optimization. This project achieved critical milestones in mixed-signal circuit stability and software-hardware co-design, directly addressing national priorities in semiconductor resiliency. The success of this work is deeply rooted in the support from the Cyclotron Road LEEP, which provided the essential "hard tech" runway—funding, mentorship, and access to Lawrence Berkeley National Laboratory’s world-class characterization facilities—allowing Vellex to overcome the "Valley of Death" often faced by deep-tech hardware startups. By leveraging LBNL’s advanced testing infrastructure, Vellex was able to rigorously benchmark the ANC architecture against state-of-the-art digital solutions, a feat that would have been resource-prohibitive independently. This collaboration has not only advanced American leadership in analog computing but has also matured Vellex’s technology to a stage ripe for private sector commercialization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Making Uintah Performance Portable for Department of Energy Exascale Testbeds

To help ease ports to forthcoming Department of Energy (DOE) exascale systems, testbeds have been made available to select users. These testbeds are helpful for preparing codes to run on the same hardware and similar software as in their respective exascale systems. This paper describes how the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system, has been modified to be performance portable across the DOE Crusher, DOE Polaris, and DOE Sunspot testbeds in preparation for portable simulations across the exascale DOE Frontier and DOE Aurora systems. The Crusher, Polaris, and Sunspot testbeds feature the AMD MI250X, NVIDIA A100, and Intel PVC GPUs, respectively. This performance portability has been made possible by extending Uintah’s intermediate portability layer [18] to additionally support the Kokkos::HIP, Kokkos::OpenMPTarget, and Kokkos::SYCL back-ends. This paper also describes notable updates to Uintah’s support for Kokkos, which were required to make this extension possible. Results are shown for a challenging radiative heat transfer calculation, central to the University of Utah’s predictive boiler simulations. These results demonstrate single-source portability across AMD-, NVIDIA-, and Intel-based GPUs using various Kokkos back-ends.

Holmen, John↗

Characterization and thermometry of dissipatively stabilized steady states

In this work we study the properties of dissipatively stabilized steady states of noisy quantum algorithms, exploring the extent to which they can be well approximated as thermal distributions, and proposing methods to extract the effective temperature T. We study an algorithm called the relaxational quantum eigensolver (RQE), which is one of a family of algorithms that attempt to find ground states and balance error in noisy quantum devices. In RQE, we weakly couple a second register of auxiliary ‘shadow’ qubits to the primary system in Trotterized evolution, thus engineering an approximate zero-temperature bath by periodically resetting the auxiliary qubits during the algorithm’s runtime. Balancing the infinite temperature bath of random gate error, RQE returns states with an average energy equal to a constant fraction of the ground state. We probe the steady states of this algorithm for a range of base error rates, using several methods for estimating both T and deviations from thermal behavior. In particular, we both confirm that the steady states of these systems are often well-approximated by thermal distributions, and show that the same resources used for cooling can be adopted for thermometry, yielding a fairly reliable measure of the temperature. These methods could be readily implemented in near-term quantum hardware, and for stabilizing and probing Hamiltonians where simulating approximate thermal states is hard for classical computers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Quantum error mitigation for Fourier moment computation

Hamiltonian moments in Fourier space—expectation values of the unitary evolution operator under a Hamiltonian at different times—provide a convenient framework to understand quantum systems. They offer insights into the energy distribution, higher-order dynamics, response functions, correlation information, and physical properties. This paper focuses on the computation of Fourier moments within the context of a nuclear effective field theory on superconducting quantum hardware. The study integrates echo verification and noise renormalization into Hadamard tests using control reversal gates. These techniques, combined with purification and error suppression methods, effectively address quantum hardware decoherence. The analysis, conducted using noise models, reveals a significant reduction in noise strength by two orders of magnitude. Moreover, quantum circuits involving up to 266 gates over five qubits demonstrate high accuracy under these methodologies when run on IBM superconducting quantum devices. Published by the American Physical Society 2025

Kiss, Oriel (ORCID:0000000174613342)↗

Miniaturized Magnetoelastic Sensor System

This article describes the design, assembly, and implementation of a hand-held, magnetic-field-based sensor system that can be adapted for a variety of sensing applications. The miniaturized system is based on Chemical Identification by Magneto-Elastic Sensing (ChIMES) technology, which uses three concentric solenoid coils to wirelessly interrogate a sensor body comprised of a response material coupled to a magnetoelastic wire. The response material expands when it encounters a target, imposing mechanical stress on the wire and altering its magnetic permeability. The sensor bodies are passive, requiring no external power source, and they are small, measuring about 15 mm in length and 3.0 mm in diameter. Up to four sensor bodies can be configured as an evenly-spaced linear array. The sensor system operates by applying a low-frequency, current-stabilized, filtered triangle wave to a uniform-density excitation coil to switch the magnetic domains within the wire. Further, the responses from the sensors are picked up by a detection coil as stress-induced changes in the Faraday voltage, and the strong magnetic field induced by the excitation coil in the detection coil is nullified by a cancellation coil reverse-wound in series with the detection coil. The responses of the sensors in an array are separated in time by a linear gradient dc biasing coil. The sensors can be interrogated through metallic and nonmetallic barriers. The signals from the detection coil and the excitation coil are digitized by a pair of bipolar analog-to-digital converters (ADCs). A Raspberry Pi single-board computer (SBC) and associated software perform data acquisition and control all aspects of the sensor system hardware. The program allows the user to select the number of sensors in the array, the type of signal that is being collected, and the number of samples to take. The program also allows for signal processing of the sensor data, such as baseline correction. The program can differentiate sensor peaks from each other and calculate the magnitude of each sensor response with less than 1% error. The data are then displayed along with a graph of the signal.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Zero and Finite Temperature Quantum Simulations Powered by Quantum Magic

We introduce a quantum information theory-inspired method to improve the characterization of many-body Hamiltonians on near-term quantum devices. We design a new class of similarity transformations that, when applied as a preprocessing step, can substantially simplify a Hamiltonian for subsequent analysis on quantum hardware. By design, these transformations can be identified and applied efficiently using purely classical resources. In practice, these transformations allow us to shorten requisite physical circuit-depths, overcoming constraints imposed by imperfect near-term hardware. Importantly, the quality of our transformations is t u n a b l e : we define a 'ladder' of transformations that yields increasingly simple Hamiltonians at the cost of more classical computation. Using quantum chemistry as a benchmark application, we demonstrate that our protocol leads to significant performance improvements for zero and finite temperature free energy calculations on both digital and analog quantum hardware. Specifically, our energy estimates not only outperform traditional Hartree-Fock solutions, but this performance gap also consistently widens as we tune up the quality of our transformations. In short, our quantum information-based approach opens promising new pathways to realizing useful and feasible quantum chemistry algorithms on near-term hardware.

Physics↗

Smart Hydro: AI Applications

This presentation provides an overview of artificial intelligence (AI) applications in hydropower.

13 HYDRO ENERGY↗

DS-TIDE: Harnessing Dynamical Systems for Efficient Time-Independent Differential Equation Solving

Time-Independent Differential Equations (TIDEs) are central to modeling equilibrium behavior across a wide range of scientific and engineering domains, from electrostatics to porous media flow. Conventional numerical solvers offer reliable solutions but incur significant computational costs due to fine-grained discretization and iterative procedures. Machine learning-based approaches address this by replacing iterative solving processes with one-time inference; however, their sophisticated models require extensive training resources that often exceed those of traditional solvers. Consequently, designing a TIDE solver that achieves high accuracy, broad applicability, and exceptional computational efficiency remains a fundamental challenge. In this paper, we propose DS-TIDE, a novel hardware solver that is inspired by, and subsequently leverages, the intrinsic connection between Dynamical Systems (DS) and Differential Equations (DEs) to efficiently and accurately solve TIDEs. DS-TIDE employs a CMOS-compatible DS-based processor, whose physical states evolve under carefully designed DE-driven dynamics and naturally converge to equilibrium -- the solution of the target TIDE -- within ~1µs on a ~1-watt DS-TIDE processor. To enhance expressivity, DS-TIDE incorporates Heterogeneous Dynamics with Temporal Layering (HDTL), which solves TIDEs through a three-stage DS evolution -- conditioning, solving, and decoding -- each governed by specialized dynamics. The entire evolution process is analogous to an infinitely deep neural network temporally unrolled, offering the system the capability of representing complex equations. Furthermore, DS-TIDE is equipped with an on-device DS-DE Auto-Alignment mechanism that dynamically adapts intrinsic hardware dynamics within milliseconds, effectively aligning the system’s dynamics to diverse target DEs. Experimental results across TIDEs from a wide range of scientific and engineering domains demonstrate that DS-TIDE achieves ~10^3× speedup, ~10^5× energy savings, and competitive or superior accuracy compared to state-of-the-art numerical and ML-based solvers.

Liu, Chuan↗

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Designing FAIR Workflows at OLCF: Building Scalable and Reusable Ecosystems for HPC Science

High Performance Computing (HPC) centers, such as the Oak Ridge Leadership Computing Facility (OLCF), provide advanced infrastructure that enables scientific research at extreme scale. These centers operate with unique hardware configurations, specialized software environments, and elevated security re quirements that differ substantially from what most users encounter on their local systems. As a result, users often develop customized digital artifacts that are tightly coupled to the specific configuration of a given HPC center. Although necessary, this practice can lead to significant duplication of effort as multiple users independently create similar solutions to common problems.

97 MATHEMATICS AND COMPUTING↗

ESnet-JLab FPGA Accelerated Transport (control plane) [EJFAT (udplbd2)] v2.0

The ESnet-JLab FPGA Accelerated Transport system is a solution for streaming high-speed scientific measurement data from Data Acquisition Systems (DAQs) to high-performance computing facilties. It is generally compatible with many science workflows, and makes no assumptions about the specifics of any particular experiment. This program (udplbd version 2) implements the control plane for the system. It is responsible for programming network forwarding rules into the data plane (implemented by the hardware designed named udplb, described separately). It also implements the control loop necessary to match up offered workload with available capacity on high-performance compute nodes.

Howard, Derek [Lawrence Berkeley National Laborato↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗

Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing

Quantum computing technology holds substantial promise as a reliable computational paradigm. However, current noisy intermediate scale quantum (NISQ) systems, are significantly impacted by noise originating from hardware inconsistencies. This noise causes errors and lowers output fidelity. So we must find which basis states cause errors. However, there are two main challenges in analyzing noise corresponding to basis states. First, the noise distribution data is high dimensional in nature, thereby making its analysis challenging. Second, although functional box plots have been used in the state of the art research to understand such a high dimensional data, they suffer from clutter and occlusion issues because of overplotting. In this study, we introduce an innovative visualization pipeline to address the aforementioned challenges to provide a clear depiction of noisy and less-noisy basis states. Specifically, our proposed visualization pipeline comprises three stages namely, low dimensional embedding, clustering, and violin plot visualization, to reduce visual clutter and effectively analyze high-dimensional noise distribution data. Our analysis uses quantum machine learning (QML) circuits as case study for drawing a distinction between noisy and less noisy basis states.

Senapati, Priyabrata [Kent State University]↗

Design, Analysis, and Experimental Testing of Hydrogen Lean Direct Injection Nozzles at Elevated Pressure

Abstract There are many challenges of commissioning a hydrogen combustor into future gas turbine engines; especially regarding achieving emissions goals. Previously, Escudero et al. and Tran et al. conducted a study to adapt the liquid fuel Lean Direct Injection (LDI) concept from Jet-A to gaseous natural gas-hydrogen blends and pure hydrogen [1], [2]. Experimental data was collected at atmospheric conditions using a Box Behnken design of experiments. The design of experiments suggested that biasing the air split in favor of the inner air circuit and increasing the swirl strength of this inner air passage resulted in improved NOx emissions, while the inverse was true for stability, which was quantified by studying the lean blowoff point (LBO) [1], [2]. The trends revealed by the original experiment [1], [2] provided a design direction for further iterations of the experimental hardware. The study presented herein describes the further investigation of such LDI injectors through experimental methods and computational fluid dynamic (CFD) simulations at atmospheric conditions, which were used to identify potential flow behaviors driving enhanced emissions performance. Further evaluation of select injectors from both studies was then conducted at elevated pressures up to 6 atmospheres. The results from both experiments are presented in this study, which include flame observations, emissions measurements, and operational challenges. NOx emissions results are reported on a volume basis in ppmvd corrected to 15% O2 and corrected for fuel. A predictive model for relating NOx emissions to test conditions at atmospheric conditions show high significance to adiabatic flame temperature while little to no significance to fuel composition for the best performing configurations. The results illustrate the connection between atmospheric testing and testing elevated pressures. The design direction indicated by the initial tests and CFD results in promising configurations for implementation into a Multi-point LDI array.

08 HYDROGEN↗