Search NASA⌕ Search

SEARCH · Search NASA

Results for “Computer Hardware”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU Systems

Deep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation. The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability. With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy. Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference. However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies. While an alternative solution is to test on GPU simulators, they are often too slow for these l

Li, Ying [William & Mary, Williamsburg, VA, USA] (↗

OSCAR-Mike: Why not ATAK?

The Office of Nuclear Smuggling Detection and Deterrence (NSDD) is evaluating methods to increase the probability that international partners will detect radioactive materials. One proposed method to accomplish this goal is to improve communications within teams operating in challenging environments. The goal of these improvements would be to enable remote monitoring of detection equipment, share data among end users in the field and subject matter experts, and integrate data from radiation detectors with other types of sensors and camera systems. Accomplishing this goal has the potential to improve the capabilities of currently deployed NSDD equipment. During FY 2021, Oak Ridge National Laboratory demonstrated some core and expanded capabilities of the Android Team Awareness Kit (ATAK), a situational awareness application developed by the US Department of Defense, to enable precision targeting, navigation, and data sharing. During FY 2022, Oak Ridge National Laboratory developed software requirements for an NSDD team awareness kit–based system. The requirements were developed by liaising with NSDD management and subject matter experts to determine NSDD’s needs and by reviewing available software and hardware solutions with US Department of Defense team awareness kit program managers and developers and radiation detection equipment vendors.

97 MATHEMATICS AND COMPUTING↗

Opportunities for retrieval and tool augmented large language models in scientific facilities

Upgrades to advanced scientific user facilities such as next-generation x-ray light sources, nanoscience centers, and neutron facilities are revolutionizing our understanding of materials across the spectrum of the physical sciences, from life sciences to microelectronics. However, these facility and instrument upgrades come with a significant increase in complexity. Driven by more exacting scientific needs, instruments and experiments become more intricate each year. This increased operational complexity makes it ever more challenging for domain scientists to design experiments that effectively leverage the capabilities of and operate on these advanced instruments. Large language models (LLMs) can perform complex information retrieval, assist in knowledge-intensive tasks across applications, and provide guidance on tool usage. Using x-ray light sources, leadership computing, and nanoscience centers as representative examples, we describe preliminary experiments with a Context-Aware Language Model for Science (CALMS) to assist scientists with instrument operations and complex experimentation. With the ability to retrieve relevant information from facility documentation, CALMS can answer simple questions on scientific capabilities and other operational procedures. With the ability to interface with software tools and experimental hardware, CALMS can conversationally operate scientific instruments. By making information more accessible and acting on user needs, LLMs could expand and diversify scientific facilities’ users and accelerate scientific output.

97 MATHEMATICS AND COMPUTING↗

Implementing a Laser Stabilization System for Trapping Ca+ Ions: an Internship Reflection

At Lawrence Livermore National Laboratory, I contributed to a project developing 3D printed micro ion traps for quantum computing. I designed, implemented, and assessed a laser stabilization system that locked lasers to the frequencies required for calibrating our High Finesse WS8-10 wavelength meter and for laser cooling and trapping of Ca+ ions. I also programmed a Python interface for hardware communication, data collection, and statistical analysis. Additionally, I optimized and aligned laser beam paths, and I implemented a closed digital feedback loop using Proportional, Integral, and Derivative (PID) control parameters. I analyzed both the long-term and short-term behavior of our locked lasers and adjusted PID parameters to enhance performance. Furthermore, I used COMSOL to simulate the capacitance of a linear Paul trap design and predict our trap’s performance. The procedures I developed for the interface, analysis, and simulations will continue to support the ion trapping experiment after my appointment. I strengthened my skills in data analysis, Python coding, and optical alignment for laser systems. My confidence as a researcher grew, particularly in communicating my research. This experience taught me the importance of careful planning and consideration in research and solidified my desire to continue exploring novel quantum technology as an undergraduate

42 ENGINEERING↗

MIND-MAC: Multi-Level In-memory Quasi Non-Destructive MAC Operation in Compact 2T-nC FeRAM for Efficient DNN Accelerator

We present MIND-MAC, a compact 2T-nC FeRAM architecture that performs multi-level, quasi-non-destructive in-memory multiply–accumulate (MAC) for deep neural networks. By exploiting voltage-controlled partial domain switching in MFM capacitors and read-transistor amplification, the cell stores multi-bit weights and gates bit-serial inputs to produce an accumulated current on shared lines. We combine TCAD-extracted parasitics with experimentally calibrated ferroelectric models in SPICE to validate device-/circuit-level behavior, and validate multi-level sensing and QNRO with measurements on a fabricated 2T-3C test vehicle. An analytical system model maps MIND-MAC to a 6-GB main-memory in-memory compute (IMC) architecture and benchmarks VGG13 inference in 61.08 ms at 964.99 mJ. Results indicate high density, reduced rewrite overhead, and energy efficiency, positioning 2T-nC FeRAM as a promising IMC candidate for next-generation AI hardware.

36 MATERIALS SCIENCE↗

SmartFuse: Reconfigurable Smart Switches to Accelerate Fused Collectives in HPC Applications

Communication switches have sometimes been augmented to process collectives (e.g., the IBM BlueGene project and the Mellanox SHArP switch). In this work, we find that there is a great acceleration opportunity through the further augmentation of switches to accelerate more complex functions that combine communication with computation. We consider three types of such functions. The first is fully-fused collectives built by fusing multiple existing collectives like Allreduce with Alltoall. The second is semi-fused collectives built by combining a collective with another computation. The third we refer to as higher-order collectives built by combining multiple computations and communications, such as to perform matrix-matrix multiply (PGEMM). In this work, we propose a framework called SmartFuse to accelerate fused collective functions. The core of SmartFuse is a reconfigurable smart switch to support these operations. The semi/fully fused collectives are implemented with a CGRAlike architecture, while higher-order collectives are implemented with a more specialized computational unit that can also schedule communication. Supporting our framework is software to evaluate and translate relevant parts of the input program, compile them into a control data flow graph, and then map this graph to the switch hardware. The proposed framework, once deployed, has the strong potential to accelerate existing HPC applications transparently by encapsulation within an MPI implementation. Experimental results show that this approach improves the performance of the PGEMM kernel, MINIFE, and AMG by, on average, 94%, 15%, and 13%, respectively.

Haghi, Pouya↗

SQMS science advances impact on Rigetti commercial processors

The collaboration between the Superconducting Quantum Materials and Systems Center (SQMS) and Rigetti Computing produced several advancements in our understanding of the role of materials characteristics in quantum processor performance. This partnership leverages SQMS's extensive characterization infrastructure and cutting-edge research in materials, and Rigetti's expertise in quantum hardware and robust nanofabrication to improve precision and performance of Rigetti's test QPUs. Qubit frequency is determined in large part by the properties of Josephson junctions (JJs) made of amorphous oxide tunnel barriers; the Alternating-Bias Assisted Annealing (ABAA) process allows us to tune JJs to their desired frequency [1]. Work by SQMS researchers in characterizing high-precision JJs post-processed (using ABAA) have yielded crucial information on the nature of the structure and chemical bonding uniformity of the ABAA processed amorphous oxides. Performance has also been improved through a comprehensive series of experiments that tested encapsulation and surface treatment. Encapsulation of the niobium metal layer with tantalum resulted in an T1 improvement of 80%, experimentally confirming the role of Nb surface losses in qubit performance [2]. Pre-treatment of the underlying silicon surface prior to JJ fabrication by replacing a buffered oxide etch (BOE) with hydrofluoric acid (HF) followed by aqueous ammonium fluoride (NH4F) has shown a statistically significant improvement of T1 by 22%, and reduction in the number of strongly-coupled TLS [3]. These examples, as well as many other published and ongoing investigations, demonstrate the mutual benefits that come from Rigetti's involvement in the SQMS collaboration. [1] - Pappas, D.P., et al. (2024). https://doi.org/10.1038/s43246-024-00596-z [2] - Bal, M., et al. (2024). https://doi.org/10.1038/s41534-024-00840-x [3] Kopas, C. J. et al. Preprint at https://doi.org/10.48550/arXiv.2408.02863 (2024).

Lachman, Ella↗

Real-Time Lifetime Prediction of Semiconductor Devices Using Hardware-in-the-Loop

This paper presents a unique approach to enable real-time lifespan prediction of semiconductor power modules using a Hardware-in-the-Loop (HIL) system. By integrating the module's overall loss characteristics-specifically switching and conduction losses-with a thermoelectric model of the thermal management system, this research demonstrates that the model can dynamically estimates the junction temperature profile of the semiconductor devices in response to a changing torque demand profile for the motor drive system. This capability enables continuous monitoring of the module's operational time and cumulative stress induced on the devices to compute accumulated remaining lifetime or time-to-failure (TTF). This study provides an architectural framework for the HIL system with high-fidelity component models of multiple physical domains, allowing simulation of dynamic behaviors of a closely-coupled motor drive system. The advanced real-time computation and measurement functionalities of the HIL system allow for both dynamic lifetime calculations based on simulated data and aggregate lifetime predictions utilizing historical data. Moreover, this paper details an algorithm that not only computes cumulative damage but also synthesizes these data into a comprehensive aggregated lifetime metric. This methodology can enhance the maintenance scheduling strategies and operational reliability of semiconductor devices in critical applications, ultimately extending their service life while optimizing performance.

hardware-in-the-loop (HIL)↗

Impact of dynamics, entanglement and Markovian noise on the fidelity of few-qubit digital quantum simulation

Quantum algorithms have been proposed to accelerate the simulation of the chaotic dynamical systems that are ubiquitous in the physics of plasmas. Quantum computers without error correction might even use noise to their advantage to calculate the Lyapunov exponent by measuring the Loschmidt echo fidelity decay rate. For the first time, digital Hamiltonian simulations of the quantum sawtooth map, performed on the IBM-Q quantum hardware platform, show that the fidelity decay rate of a digital quantum simulation increases during the transition from dynamical localization to chaotic diffusion in the map. The observed error per CNOT gate increases by $1.5{\times }$ as the dynamics varies from localized to diffusive, while only changing the phases of virtual RZ gates and keeping the overall gate count constant. A gate-based Lindblad noise model that captures the effective change in relaxation and dephasing errors during gate operation qualitatively explains the effect of dynamics on fidelity as being due to the localization and entanglement of the states created. Specifically, highly delocalized states that are entangled with random phases show an increased sensitivity to dephasing and, on average, a similar sensitivity to relaxation as localized states. In contrast, delocalized unentangled states show an increased sensitivity to dephasing but a lower sensitivity to relaxation. This gate-based Lindblad model is shown to be a useful benchmarking tool by estimating the effective Lindblad coherence times during CNOT gates and finding a consistent $2\unicode{x2013}3{\times }$ shorter $T_2$ time than reported for idle qubits. Thus, the interplay of the dynamics of a simulation with the noise processes that are active can strongly influence the overall fidelity decay rate.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Flow dynamics and heat transfer in simplified battery energy storage systems with heated battery modules

Large-scale energy storage systems (ESSs) composed of batteries show promise in addressing current energy challenges, but dissipation of generated heat is important. Here, this paper focuses on buoyant convective flows in simplified ESS battery racks. Natural convection is not generally the primary cooling strategy but can be important in abnormal scenarios where there is module overheat or potentially thermal runaway. We use computational fluid dynamics to investigate the flow dynamics and heat transfer mechanisms in a simplified parameterized rack design. Despite its simplicity, this configuration produces many of the relevant features expected in real ESSs without details of module geometry or hardware, allowing broad conclusions independent of manufacture-specific designs. We start by providing visualizations of the flowfield and measurements of entrainment, heat flux, and pressure. To characterize the dependence on the system parameters, we develop an integral-scale analysis of the average temperature equation to highlight the dominant source terms. We use results from this analysis to derive a steady network model composed of simple algebraic expressions to provide first-order predictions of entrainment through the rack. The network model leads to a linear scaling of the Reynolds number based on convective mass flux with respect to the Grashof number based on the heat source. We deduce empirical relationships that relate the heat exchanged between modules using a surface-averaged Nusselt number as a function of the local Reynolds and Rayleigh numbers. Lastly, we investigate how space between the modules and rack in the spanwise direction creates flow bypass, resulting in different flow pathways.

Battery thermal management↗

Towards utility-scale electronic structure with sample-based quantum bootstrap embedding

One of the main applications for which quantum computers are hoped to find utility is in simulating ground state energies and other observables of molecular chemical systems. The recently proposed sample-based diagonalization method is a readily implementable method for this task on current-day hardware using short circuit depths and has been demonstrated on as many as 85 qubits in recent studies. In this work, we combine the recently proposed quantum bootstrap embedding (QBE) method with sampled-based diagonalization (QBE-SQD) and present the first benchmarking study of the QBE method on real quantum hardware, ibm_pittsburgh, a Heron r3 processor with 156 qubits. Our test system is a hydrogen ring with 8 hydrogen atoms in the cc-pVDZ basis. We show that for this system, QBE-SQD using an active space of (8e, 19o) per fragment with a 43 qubit footprint produces a ground state energy accuracy which exceeds that of an SQD calculation with an (8e, 30o) active space with a 67 qubit footprint when using a comparable number of Slater determinants. This demonstrates that the use of quantum bootstrap embedding techniques is a promising path towards extending the capabilities of state-of-the-art quantum eigensolvers on near-term devices.

Bierman, Joel [North Carolina State University, Ra↗

Temperature and Pressure Instrumentation for LYNM PE1 Chemical Explosive Testing

Underground chemical explosive testing has been conducted at the Nevada National Security Site under the Physics Experiment 1 (PE1) to validate explosive computer modeling and, ultimately, improve the accuracy of subsurface explosive detection. This SAND Report describes the dynamic temperature and pressure measurements within the chamber induced by the chemical explosive for the first of three experiments, PE1-A. The report details the instrumentation used for the experiment, the emplacement of the hardware, and the measured results. Dynamic temperature measurements were accomplished with the use of optical spectrometers and dynamic pressure was measured with a series of high-rated pressure transducers. This report includes details of the design and results of four cavity sensor systems used to measure early-time temperature, early-time pressure, late-time temperature, and late time pressure. The outcomes of PE1-A were used to inform the design of the remaining PE1 series experiments, PE1-B and PE1-DL.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Fully quantum algorithm for mesoscale fluid simulations with application to partial differential equations

Fluid flow simulations marshal our most powerful computational resources. In many cases, even this is not enough. Quantum computers provide an opportunity to speed up traditional algorithms for flow simulations. We show that lattice-based mesoscale numerical methods can be executed as efficient quantum algorithms due to their statistical features. This approach revises a quantum algorithm for lattice gas automata to reduce classical computations and state preparation at every time step. For this, the algorithm approximates the qubit relative phases and subtracts them at the end of each time step. Phases are evaluated using the iterative phase estimation algorithm and subtracted using single-qubit rotation phase gates. Further, this method optimizes the quantum resource required and makes it more appropriate for near-term quantum hardware. We also demonstrate how the checkerboard deficiency that the D1Q2 scheme presents can be resolved using the D1Q3 scheme. The algorithm is validated by simulating two canonical partial differential equations: the diffusion and Burgers' equations on different quantum simulators. We find good agreement between quantum simulations and classical solutions for the presented algorithm.

97 MATHEMATICS AND COMPUTING↗

EUREICA: Efficient UltRa Endpoint IoT-enabled Coordinated Architecture

The electricity grid has evolved from a physical system to a cyber-physical system with digital devices that perform measurement, control, communication, computation, and actuation. The increased penetration of distributed energy resources (DERs) that include renewable generation, flexible loads, and storage provides extraordinary opportunities for improvements in efficiency and sustainability. However, they can introduce new vulnerabilities in the form of cyberattacks, which can cause significant challenges in ensuring grid resilience. The purpose of this project was to develop a framework ((Efficient, Ultra-REsilient, IoT-Coordinated Assets, or EUREICA)for achieving grid resilience through suitably coordinated assets including a network of Internet of Things (IoT) devices, and a local electricity market (LEM) to identify trustable assets and carry out this coordination. Situational Awareness (SA) of locally available DERs with the ability to inject power or reduce consumption is enabled by the market, together with a monitoring procedure for their trustability and commitment. Experiments conducted during this project demonstrated that, with this SA, a variety of cyberattacks can be mitigated using local trustable resources without stressing the bulk grid. The demonstrations were carried out using a variety of high-fidelity co-simulation platforms, real-time hardware-in-the-loop validation, and a utility-friendly simulator.

14 SOLAR ENERGY↗

ARCH Technology Snapshot Autonomous Robot Control Hierarchy (ARCH): A universal software system that removes the need to rebuild robotic software for every new platform or task

Robots are increasingly used to perform repetitive, hazardous, and time-sensitive tasks, improving safety and operational efficiency. However, most robotic systems remain difficult to adapt because they are tightly tied to specific hardware and require extensive reprogramming for each new configuration.

42 ENGINEERING↗

Combining Deep Learning and scatterControl for High-Throughput X-ray CT Based Non-Destructive Characterization of Large-Scale Casted Metallic Components

X-ray computed tomography (XCT) is essential for nondestructive evaluation and quality control of large-scale metal components. XCT imaging, however, faces significant challenges from metal artifacts, particularly those caused by Compton scattering, which degrade image quality and obscure critical details. Hardware-based solutions (e.g. scatterControl) offer advancements by intercepting scattered photons and reducing artifacts, but they can be time-consuming and require additional processing. Here, we propose modifying and leveraging a novel deep learning (DL) framework, Simurgh, to enhance and accelerate scatter correction in XCT. By combining scatterControl with DL-based artifact removal, we demonstrate significant reduction in scan time while producing high-quality reconstructions. Through extensive evaluation on industrial XCT data, we show that our methods reduce scan time by up to more than 10 x while preserving flaw detectability. Quantitative analysis across multiple segmentation techniques confirms that Simurgh-based reconstructions consistently outperform traditional Feldkamp-Davis-Kress, model-based iterative reconstruction, and commercial DL models in both pixel-level and task-specific evaluations, enabling scalable, high-throughput XCT workflows for characterization of large scale components in applications such as casting and metal additive manufacturing.

Complex metal parts↗

EJFAT Scientific Perspective

Presented new computing model to the test by deploying the EJFAT system alongside a data-stream processing framework running the production-level CLAS12 event reconstruction application. In this experiment, a continuous stream of CLAS12 Level-1 identified events was processed in real-time using the EJFAT load balancer, distributing the workload across 90 computing nodes located across the U.S. This marks the first-ever large-scale, real-time distributed data stream processing experiment, demonstrating that scientific data-streaming pipelines can efficiently scale across four dimensions, thanks to EJFAT’s advanced hardware and software capabilities.

Gyurjyan, Vardan [Thomas Jefferson National Accele↗

Toward computing bounds for Ramsey numbers using quantum annealing

Quantum annealing is a powerful tool for solving and approximating combinatorial optimization problems, such as graph partitioning, community detection, centrality, routing problems, and more. In this paper we explore the use of quantum annealing as a tool for use in exploring combinatorial mathematics research problems. We consider the monochromatic triangle problem and the Ramsey number problem, both examples of graph coloring. Conversion to quadratic unconstrained binary optimization (QUBO) form is required to run on quantum hardware. While the monochromatic triangle problem is quadratic by nature, the Ramsey number problem requires the use of order reduction methods for a quadratic formulation. The goal is to provide a method for producing special colorings of graphs which if successful would provide lower bounds for certain Ramsey numbers. We discuss implementations, limitations, and results when running on the D-Wave Advantage quantum annealer.

97 MATHEMATICS AND COMPUTING↗