Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Deep Learning Super-Resolution X-Ray Computed Tomography Algorithms for Additive Manufacturing

Industrial X-ray computed tomography (XCT) is a nondestructive method for inspection and characterization of additively manufactured (AM) materials and parts. In practice, the resolution of XCT can be limited by factors such as detector binning, restricted field of view for large-scale objects, system blur, motion during scanning, and acquisition settings. These limitations can reduce the detectability of critical flaws such as pores, cracks, and lack of fusion. Super-resolution (SR) techniques offer a promising solution for improving the effective resolution and image quality of XCT reconstructions without the need for expensive hardware upgrades or laborious, time-consuming scans. In particular, deep learning-based SR methods have garnered attention in recent years as powerful tools for reconstructing high-resolution volumes from low-resolution inputs. In this work, a novel deep learning-based SR method is proposed for XCT scans of AM parts, and compared against several existing state-of-the-art (SOTA) methods. The proposed method, Simurgh-SR, is built on the pre-existing Simurgh framework and consists of a 2.5D U-Net trained to map low-quality inputs containing noise and artifacts to high-quality reconstructions characterized by higher flaw contrast, better noise texture, and reduced artifacts. The experimental results demonstrate superior performance of Simurgh-SR in performing 4× SR on real industrial XCT scans of thick 316L components, enhancing the structural similarity score and peak signal-to-noise ratio (>7dB) compared to the LR counterpart while improving the F1-score for flaw detection by more than 2.3× when compared to alternative SOTA SR methods. This improvement enables more accurate and significantly faster characterization of metal AM components. Additionally, Simurgh-SR was trained for both 2X and 4X SR and performs effectively at both levels, enabling the use of a single model for various SR factors.

Rahman, Obaid [ORNL] (ORCID:0000000277810840)↗

Evaluating Variable-Impedance Magnetically-Insulated Transmission Lines as a Risk-Mitigation Measure for Next-Generation Pulsed Power

This project has produced the first detailed characterizations of power flow resulting from applying the “variable-impedance MITL” concept to real-life systems in Sandia’s pulsed power program (Z and next-generation pulsed power (NGPP)). We present simulation results and analyses for constant-impedance versions of both Z and NGPP and survey the operational viability of several variable-impedance re-designs in the parameter space of linear tapers. Circuit modeling (SCREAMER/Bertha) was used to pinpoint promising candidate designs, and EM-PIC (Empire) simulations were used to evaluate these candidates more rigorously. This approach was particularly successful in the Z regime which resulted in the identification of several viable variable-impedance MITL designs for each level. The approach was more challenged in the operating space NGPP occupies, producing data points that speak to a more restrictive design space due to anode plasma turn-on. In the end, we were able to converge on one viable variable-impedance design for the highest inductance line (level “F”) and one for the highest current line (level “A”). Altogether, the body of simulation evidence presented in this report suggest there does exist flexibility in operating space for magnetically-insulated transmission lines (MITLs) having variable geometric impedance to be a potential enabling technology for safely increasing current delivery (and potentially lowering stack voltage) in pulsed-power drivers by manipulating electron losses; however, operating points for a particular design must be carefully screened. Circuit and EM-PIC modeling provided consistent verdicts in safe operating regimes for operational viability, but additional physics such as anode plasma turn-on which is included in Empire but not in SCREAMER/Bertha was found to be a critical factor affecting power flow that lead to different assessments between the codes. It is not always the case that the occurrence of anode plasma caused a design to fail (some designs turned on anode plasma yet still delivered load currents meeting design targets); the details matter such as how early in the pulse anode surfaces break down (and how large a region). However, in every case that it did fail it was found that the feedback from anode plasma was the cause (i.e., turning off the anode plasma model in Empire restored agreement with the circuit model prediction). As circuit simulations represent an efficient and practical means of surveying design space compared to more computationally-expensive approaches such as EM-PIC, it could be prudent to invest in the research and development of models to include the effects of anode plasma such as ion emission in circuit codes. The variable-impedance MITL design is a new concept that enables controlled manipulation of the initial electron losses in the outer MITL and can be tested on Z today. We encourage follow-on work to explore further optimization (including alternative variable-impedance profiles, e.g., having constant dZ/dR), and to confirm the major findings presented in this report by fielding test hardware on actual Z shots.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Autonomous elemental characterization enabled by a low cost robotic platform built upon a generalized software architecture

Despite the rapidly growing applications of robots in industry, the use of robots to automate tasks in scientific laboratories is less prolific due to the lack of generalized methodologies and the high cost of hardware. This paper focuses on the automation of characterization tasks necessary for reducing cost while maintaining generalization and proposes a software architecture for building robotic systems in scientific laboratory environments. A dual-layer (Socket.IO and ROS) action server design is the basic building block, which facilitates the implementation of a web-based front end for user-friendly operation and the use of ROS Behavior Trees for convenient task planning and execution. A robotic platform for automating mineral and material sample characterization is built upon the architecture, with an open-source, low-cost three-axis computer numerical control gantry system serving as the main robot. A handheld laser induced breakdown spectroscopy (LIBS) analyzer is integrated with a 3D printed adapter, enabling (1) automated 2D chemical mapping and (2) autonomous sample measurement (with the support of an RGB-Depth camera). We demonstrate the utility of automated chemical mapping by scanning the surface of a spodumene-bearing pegmatite core sample with a 1071-point dense hyperspectral map acquired at a rate of 1520 bits per second. Furthermore, we showcase the autonomy of the platform in terms of perception, dynamic decision-making, and execution, through a case study of LIBS measurement of multiple mineral samples. The platform enables controlled and autonomous chemical quantification in the laboratory that complements field-based measurements acquired with the same handheld device, linking resource exploration and processing steps in the supply chain for lithium-based battery materials.

Cao, Xuan [Lawrence Berkeley National Laboratory (↗

Obstacles to Practical Digital Supply Chain Risk Management in the Energy Sector

Cyber supply chain risk management (C-SCRM) programs must consider operations that depend on the lifecycles of digital components such as hardware, firmware, software, and services. We integrate academic literature, historical incidents, and existing standards to identify obstacles faced by C-SCRM programs.

Business Process Management & Integration↗

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES↗

EV SALaD 2023 Demonstration: Best Practices and Mitigations for Protecting EVSE Infrastructure

The Electric Vehicle Secure Architecture Laboratory Demonstration (EV SALaD) program is a demonstration of cybersecurity best practices for high-power electric vehicle (EV) charging infrastructure led by Idaho National Laboratory (INL), in collaboration with other DOE National Laboratories participating in the EVs at Scale Consortium.a Sandia National Laboratories (SNL) and Pacific Northwest National Laboratory (PNNL) participated in the first 2-year (FY22-23) demonstration cycle for EV SALaD. This report documents the FY23 demonstration, the second in a series of demonstrations and collaborations in deploying and operating cybersecure EV charging infrastructure. It includes a summary of improvements from the FY22 demonstration, technical analysis of the FY23 demonstration, how the research demonstrates cyber-physical and cybersecurity best practices for high-power EV charging infrastructure, and related impacts to national and energy security. For EV SALaD, the FY22 demonstration focused on the detection, ranking, and prioritization of anomalous events for high-power EV charging. The FY23 demonstration additionally included the demonstration of cybersecurity best practices, which included protection and mitigation solutions to prevent, respond, and recover from anomalous events. During the demonstrations, the multi-lab EV SALaD team conducted a Test Effect Payload (TEP)b evaluation on extreme fast charger (XFC) hardware equipped with Cerberus, a detection and response solution, to demonstrate anomaly detection and mitigation cybersecurity best practices against cyber-enabled events.

33 ADVANCED PROPULSION SYSTEMS↗

Magnetic hysteresis experiments performed on quantum annealers

While quantum annealers have emerged as versatile and controllable platforms for experimenting on correlated spin systems, the important phenomenology of magnetic memory and hysteresis remain unexplored on hardware designed to escape metastable states via quantum tunneling. Here, we present the first general protocol to experiment on magnetic hysteresis on programmable quantum annealers and implement it on three D-Wave superconducting qubit quantum annealers, using up to thousands of spins, for both ferromagnetic and disordered Ising models, and across different graph topologies. We observe hysteresis loops whose area depends nonmonotonically on quantum fluctuations, exhibiting both expected and unexpected features, such as disorder-induced steps and nonmonotonicities. Our work establishes quantum annealers as a platform for probing nonequilibrium emergent magnetic phenomena, thereby broadening the role of analog quantum computers into foundational questions in condensed matter physics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Hls4ml Synthesis Testing

HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

hls4ml

hls4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where hls4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

A modular GUI-based program for genetic algorithm-based feedback-assisted wavefront shaping

Abstract We have developed a modular graphical user interface (GUI)-based program for use in genetic algorithm-based feedback-assisted wavefront shaping. The program uses a class-based structure to separate out the universal modules (e.g. GUI, multithreading, optimization algorithms) and hardware-specific modules (e.g. code for different SLMs and cameras). This modular design makes the program easily adaptable to a wide range of lab equipment, while providing easy access to a GUI, multithreading, and three optimization algorithms (phase-stepping, simple genetic, and microgenetic).

97 MATHEMATICS AND COMPUTING↗

Classical and quantum simulations of 1+1-dimensional ${\mathbb{Z}}_{2}$ gauge theory at finite temperature and density

Simulating strongly coupled gauge theories at finite temperature and density is a longstanding challenge in nuclear and high-energy physics with fundamental implications for condensed matter physics. Here, we simulate such systems using minimally entangled typical thermal state (METTS) approaches, which combine classical random sampling with imaginary-time evolution, implementable on either classical or quantum computers, to estimate thermal averages of observables. We study 1+1-dimensional ${\mathbb{Z}}_{2}$ gauge theory coupled to spinless fermionic matter, which maps onto a local quantum spin chain. We benchmark both a classical matrix-product-state implementation of METTS and a recently proposed adaptive variational approach for near-term quantum devices, focusing on the equation of state and measures of fermion confinement. Of particular importance is the choice of basis for METTS sampling, which impacts both the sampling overhead and quantum circuit complexity. Our work sets the stage for future studies of strongly coupled gauge theories using classical and quantum hardware.

Chen, I-Chi [Iowa State Univ., Ames, IA (United St↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Thermonuclear Burn in a Multiphysics Code on GPUs

Multiphysics codes links to a library called SINGE for the calculation of thermonuclear (TN) burn rates, but some current multiphysics codes do not attempt to leverage the support for parallel operation that SINGE provides. Our goal is to investigate implementations of the SINGE workflow and analyze how the use of a performance portability layer could reduce run time on CPU archi tectures while also supporting GPU architectures without requiring code modifications. We looked to the Kokkos C++ Performance Portability Ecosystem to implement hardware agnostic parallel patterns.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Towards a Verifiable Domain-Specific Language for Hardware-Accelerated Stencils

Defining a domain-specific language (DSL) that supports vector-calculus abstractions eases the porting of partial differential equation (PDE) solvers to specialized architectures. Sufficiently high-level abstractions empower users to express universal laws with sufficient generality that the laws must always hold true within their domain of validity. A broad class of PDE solvers employs stencil-based algorithms, the target domain of Berkeley Lab's stencil accelerator chip co-design project. First released as open-source in January 2026, the Formal software framework lays a foundation for defining an embedded DSL based on composable operators that implement mimetic numerical methods -- stencil algorithms that guarantee satisfaction of discrete versions of important vector calculus theorems. The Formal DSL will be the frontend to a new class of stencil-PDE accelerators developed jointly by LBNL, UHCL, and UC Berkeley through the DOE Competitive Portfolios for Computer Science Project. This offers the potential of an order of magnitude acceleration for this important category of computational methods to serve the DOE mission. Future work on the Formal DSL will facilitate software verification via type-safe templates that enable problem-specific correctness proofs relying upon generic function theory and carefully crafted unit tests.

Rouson, Damian↗

Digital Twin for Chemical Science (DTCS) v0.01

Directly visualizing the trajectories of chemistry can unravel novel insights into the behavior of catalysts, gas phase reactions, photo-induced dynamics, and building blocks for quantum information processing. The ability of explicitly identifying, tracking, and tagging the exchange of matter, hence the annihilation and creation of new chemical species, can be best realized through a close coupling of theory and experiment. While the synchrotron-based characterization facilities propelled rapidly in its hardware, providing higher brightness, better resolution, and more precision, the software infrastructure is lagging. We developed DTCS (Digital Twin for Chemical Science) v.01, a central platform that faithfully mimics advanced instrumentations in Scientific User Facilities, by solving a variety of technical challenges in data acquisition, analysis, and model-driven interpretation. Rooted in physics and accelerated by AI, we validated this concept by direct comparison with precise experimental X-ray Photoelectron Spectroscopy (XPS) observations using a ubiquitous metal-water interfacial scenario, i.e., Ag/H2O as our main narrative. The DTCS v.01 input mirrors how the bench chemists work, with the output directly linked to the end station computer, thereby providing a user-friendly, knowledge-driven, and accessible user experience with mechanistic insights standardized in a way that are ready to be published, versioned, and transferred flexibly.

Qian, Jin↗

National Security Programs - Cyber: MMAREJBLIGE – Modular Multi Agent Grid Emulation for Joined Breakdowns in Linked Generative Emulations - 23-0644

Modular Multi Agent Grid Emulations for Joined Breakdowns in Linked Generative Emulations (MMAREJBLIGE) introduces an agent-based modeling framework into real-time cyber-physical emulation to achieve a context-aware environment that introduces operator/attacker/external-condition variability to improve emulation fidelity and testing rigor. We detail our agent framework design, internal communication via message passing, and time synchronization, as well as the individual components of the system. We include a brief analysis of several scenarios run on a real-time, hardware-in-the-loop, Industrial Control Systems (ICS) test-bed which include normal operation, physical disruption, disruption with mitigation, and disruption with mitigation during a cyber denial-of-service (DOS) attack.

42 ENGINEERING↗

CACTUS: Chemistry Agent Connecting Tool Usage to Science

Large language models (LLMs) have shown remarkable potential in various domains but often lack the ability to access and reason over domain-specific knowledge and tools. In this article, we introduce Chemistry Agent Connecting Tool-Usage to Science (CACTUS), an LLM-based agent that integrates existing cheminformatics tools to enable accurate and advanced reasoning and problem-solving in chemistry and molecular discovery. We evaluate the performance of CACTUS using a diverse set of open-source LLMs, including Gemma-7b, Falcon-7b, MPT-7b, Llama3-8b, and Mistral-7b, on a benchmark of thousands of chemistry questions. Our results demonstrate that CACTUS significantly outperforms baseline LLMs, with the Gemma-7b, Mistral-7b, and Llama3-8b models achieving the highest accuracy regardless of the prompting strategy used. Moreover, we explore the impact of domain-specific prompting and hardware configurations on model performance, highlighting the importance of prompt engineering and the potential for deploying smaller models on consumer-grade hardware without a significant loss in accuracy. By combining the cognitive capabilities of open-source LLMs with widely used domain-specific tools provided by RDKit, CACTUS can assist researchers in tasks such as molecular property prediction, similarity searching, and drug-likeness assessment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗