Search NASA⌕ Search

SEARCH · Search NASA

Results for “hardware design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Electric Drive Technologies Research: ELT223 Component Modeling, Co-Optimization, and Trade-Space Evaluation Annual Report

This project is intended to support the development of new traction drive systems that meet the targets of 100 kW/L for power electronics and 50 kW/L for electric machines with reliable operation to 300,000 miles. To meet these goals, new designs must be identified that make use of state-of-the-art and next-generation electronic materials and design methods. Designs must exploit synergies between components, for example converters designed for high-frequency switching using wide band gap (WBG) devices and ceramic capacitors. This project included: (1) a survey of available technologies; (2) investigating new technologies, that for example, reduce volume of thermal management or magnetic components; (3) the development of computer aided design tools that consider the converter volume, reliability, and electrical performance; (4) exercising the design software to evaluate performance gaps and predict the impact of certain technologies and design approaches, i.e. GaN semiconductors, ceramic capacitors, ceramic thermal management components, and select topologies; (5) building and testing hardware prototypes to validate models and concepts. The design tools enable co-optimization of the power module and passive elements and provide some design guidance. At the end of the project, new advanced computing methods, such as machine learning approaches, were considered.

33 ADVANCED PROPULSION SYSTEMS↗

Fast, Controllable, and Modular Solid-State Circuit Breaker Design for Battery Management Systems

Electric power grid is experiencing a growing number of distributed and inertia-free generation resources. To facilitate the growing generation and load demands, and ensure stable operation, energy storage systems, especially behind-themeter-storage (BTMS), have emerged as a potential candidate. BTMS plays a vital role in the grid storage sector and supports high power charging for EVs. However, the potential of thermal runaway and associated safety concerns in the batteries can hamper their widespread adoption. In this work, we provide a solution for a fast and controllable discharge of a cell that was identified as a stressful/faulty, in the battery pack, and fast circuit breaking leveraging the Solid-State Circuit Breaker (SSCB) technology which provides active control over the cell connection as opposed to conventional passive solutions. Testing on the simulation platform successfully validated the concept, demonstrating its efficacy. Controlled discharge testing with 20Ah LiFePO4 Lithium Iron Phosphate (LFP) cells from 1C-10C current rate on the hardware prototype corroborated simulation results, demonstrating design feasibility and providing essential data for its performance and thermal characteristics, while also revealing limitations that inform areas for further optimization.

33 ADVANCED PROPULSION SYSTEMS↗

Design and Demonstration of a NH3-Fueled Two-Stroke Uniflow Engine for Greenhouse Gas Reduction

The maritime shipping industry is growing increasingly interested in both low and non-carbon-containing fuels to meet future greenhouse gas emission targets. Specifically of interest is ammonia, as it has a relatively high volumetric energy density compared to other future fuels, such as hydrogen, making it more economical to transport. The robust engine architecture of low-speed two-stroke marine engines makes them an ideal candidate for ammonia fuel, overcoming many of the issues surrounding its poor ignitability and low flame speed. If emissions and fueling system challenges can be addressed, retrofits of current low-speed two-stroke dual-fuel engines represent a viable pathway for bringing ammonia engines to market. This study explores these technical hurdles by describing the design, analysis, and experimental validation of a single cylinder research engine converted to operate on ammonia fuel. The engine is a reduced-scale uniflow two-stroke marine engine with two previous hardware configurations available – diesel and high-pressure CNG dual-fuel. A concept study was used to evaluate possible ammonia-fueled engine architectures and the associated tradeoffs and design considerations. With the chosen architecture, low-pressure dual fuel, 1D and 3D analysis tools were used to inform hardware selection and to determine hardware configurations which minimized ammonia-slip. In addition to these considerations the hardware and engine configuration were designed to provide a versatile and robust testing platform. This includes options to test both gaseous and liquid ammonia injection, as well as a wide range of performance parameters such as AFR, swirl, valve timing, SOI, and many others. Design constraints imposed by the existing engine hardware necessitated an iterative loop between design and analysis toolsets, ultimately converging on a final design for the ammonia-conversion hardware. The engine was rebuilt with the new hardware and evaluated in an engine test cell. A new control strategy developed and flashed onto a prototyping electronic control unit allowed for full control over all engine parameters. An initial calibration was developed, providing test data for validation of the engine 1D and 3D models. The impact of the design choices on engine operability and the ability to meet program targets is discussed as well as opportunities for further optimization of the ammonia-conversion hardware, informed by the validated models.

Kaul, Brian [ORNL] (ORCID:0000000184813620)↗

A cross-platform execution engine for the quantum intermediate representation

Hybrid languages like the quantum intermediate representation (QIR) are essential for programming systems that mix quantum and conventional computing models, while execution of these programs is often deferred to a system-specific implementation. Here, we develop the QIR Execution Engine (QIR-EE) for parsing, interpreting, and executing QIR across multiple hardware platforms. QIR-EE uses LLVM to execute hybrid instructions specifying quantum programs and, by design, presents extension points that support customized runtime and hardware environments. We demonstrate an implementation that uses the XACC quantum hardware-accelerator library to dispatch prototypical quantum programs on different commercial quantum platforms and numerical simulators, and we validate execution of QIR-EE on IonQ, Quantinuum, and IBM hardware. Our results highlight the efficiency of hybrid executable architectures for handling mixed instructions, managing mixed data, and integrating with quantum computing frameworks to realize cross-platform execution.

LLVM↗

Design and performance of AI agents interfacing with an atomic layer deposition tool

In this work, we introduce the design of an atomic layer deposition (ALD) reactor augmented with an AI interface for autonomous materials synthesis. Our modular design encapsulates the particularities of the hardware behind a Python interface that communicates with the ALD control software via transmission control protocol. This interface is compatible with model context protocol interfaces used in agentic frameworks. We have integrated our tool with a simple AI agent that leverages a large language model to transform user-supplied queries into ALD processes that are then run in our reactor. Our approach uses a JavaScript object notation schema to encode ALD processes. Our experimental results show that the AI interface does not impose a significant overhead to our control software, at least within our fastest 10 ms scale. We also carried out a detailed evaluation of the agent performance using leading models in two classes of tasks: basic instruction and process discovery tasks, where the agent is presented with a target material and needs to identify the correct ALD process compatible with the reactor configuration. Despite the simplicity of our agent design, we observed that most of the advanced models excelled at the instruction tasks. However, only recent models, such as o1, o3, GPT-5, and Claude Opus 4, performed well in process discovery tasks. We also observed significant variability in the response for the hardest challenges. While the results obtained are promising, we identify areas where AI research could improve the performance of agents for ALD.

47 OTHER INSTRUMENTATION↗

Fault localization in a microfabricated surface ion trap using diamond nitrogen-vacancy center magnetometry

Here, as quantum computing hardware becomes more complex with ongoing design innovations and growing capabilities, the quantum computing community needs increasingly powerful techniques for fabrication failure root-cause analysis. This is especially true for trapped-ion quantum computing. As trapped-ion quantum computing aims to scale to thousands of ions, the electrode numbers are growing to several hundred, with likely integrated photonic components also adding to the electrical and fabrication complexity, making faults even harder to locate. In this work, we used a high-resolution quantum magnetic imaging technique, based on nitrogen-vacancy centers in diamond, to investigate short-circuit faults in an ion trap chip. We imaged currents from these short-circuit faults to ground and compared them to intentionally created faults, finding that the root cause of the faults was failures in the on-chip trench capacitors. This work, where we exploited the performance advantages of a quantum magnetic sensing technique to troubleshoot a piece of quantum computing hardware, is a unique example of the evolving synergy between emerging quantum technologies to achieve capabilities that were previously inaccessible.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Qudit Designs and Where to Find Them

Unitary t-designs are some of the most versatile tools in quantum information theory. Their applications range from randomized benchmarking and shadow tomography, to more fundamental ones such as emulating quantum chaos and establishing exponential separations between classical and quantum query complexity. While unitary designs originating from a group structure, such as the Clifford group, have proven to be incredibly useful for qubit systems, unfortunately, this is no longer true for qudits. In fact, the classification of finite-group representations rules out the existence of unitary 2-designs for arbitrary qudit dimensions. This severely limits the applicability of standard quantum information primitives when it comes to qudit systems. We overcome these limitations with a three-fold contribution. First, we introduce a general technique to construct families of weighted state t-designs in arbitrary qudit dimensions. These weighted state-designs generalize classical shadow tomography protocol from qubits to qudits. Second, we introduce a Clifford character RB that allows us to benchmark the qudit Clifford group in any dimension, including non-prime-power dimensions. And third, we establish bounds on the quantum circuit complexity of generating approximate unitary-designs from native gates in existing quantum hardware such as high-spin and cavity-QED qudits. Our work further highlights the analogy between spin and optical coherent states by proving that spin-GKP codewords form a state 2-design while spin coherent states do not; in direct analogy with the optical case. This work is structured as a pedagogical and self-contained introduction to unitary designs and their applications to qudit systems.

Anand, Namit [NASA, Ames; Unlisted, US] (ORCID:000↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Leveraging dendritic complexity for neuromorphic computing

Abstract Beyond-von Neumann computing approaches are necessary to sustain the growth of microelectronics and the increasing appetite for artificial intelligence/machine learning algorithms. Neuromorphic computing is an emerging paradigm that takes inspiration from the brain to provide a path forward to improve the computational efficiency and computational density of next-generation computing architectures. In nature, we observe brains performing complex computations with a much smaller energy footprint than conventional computing approaches. Current neuromorphic systems are focused primarily on scalability, namely, increasing the number of computational units (neurons) and connections between units (synapses). However, for brain-like cognition and efficiency in next-generation computing hardware, we need increased complexity in function, as well as improved connection density for scalability. Here, we present our work that aims to incorporate dendrites for ‘compute-on-wire’ in neuromorphic architectures to increase the computational complexity (e.g. number of programmable parameters, nonlinear dynamics) as well as computational efficiency (energy/compute) of artificial neural networks (ANNs). We do this by showcasing neuromorphic dendrite elements that can be leveraged for various applications. We will present examples of neuroscience-inspired direction-selective circuits and an ANN with active dendrites leveraging shunting inhibition. We also demonstrate the benefits of using dendrites in deep neural networks. To conclude, we discuss how we can utilize emerging hardware devices in these systems and design next-generation neuromorphic architectures with dendrites.

Cardwell, Suma G. (ORCID:0000000226575545)↗

A CHIL Validation of Machine Learning-Assisted Methods for Real-Time Controls of Solar PV for Grid Services

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been proposed; however, these technologies lack comprehensive validation under real-world application scenarios. This paper addresses this gap by designing and developing a controller-hardware-in-the-loop framework to evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. Simulation results indicate the superior performance of an ML-based approach compared to the conventional reference-control grouping-based approach, showcasing its potential to support grid stability and operational efficiency.

closed-loop validation↗

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Optical Stochastic Cooling Program at Fermilab

Recently, Optical Stochastic Cooling (OSC) became the first demonstrated method for ultra-high-bandwidth stochastic cooling. The initial experiments at Fermilab’s IOTA ring explored the essential physics of the method and demonstrated cooling, heating and manipulation of beams and single particles. Having been validated in practice, with continued development, OSC carries the potential for dramatic advances in the state-of-the-art performance and flexibility for beam cooling and control. The ongoing program at Fermilab is now focused on the development of an OSC system that includes high-gain optical amplification, which promises a two-order-of-magnitude increase in the strength of the OSC force. Here we review the progress and plans for the amplified OSC program. This includes detailed lattice designs and tracking simulations for the various experimental configurations, designs and status for the various hardware systems, and near-term operational plans and use cases.

Jarvis, J. [Fermilab]↗

A CHIL Validation of Machine Learning-Assisted Methods for Real-Time Controls of Solar PV for Grid Services

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been proposed; however, these technologies lack comprehensive validation under real-world application scenarios. This paper addresses this gap by designing and developing a controller-hardware-in-the-loop framework to evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. Simulation results indicate the superior performance of an ML-based approach compared to the conventional reference-control grouping-based approach, showcasing its potential to support grid stability and operational efficiency.

14 SOLAR ENERGY↗

Programmable Digital Devices used in Advanced Reactors

This paper introduces the concepts of common cause failure, diversity, and defense-in-depth used by the nuclear industry to analyze resilience in reactors. A survey of publicly traded and private companies building advanced reactors and their licensing status is presented. Safety and non-safety systems found in the NuScale Power design are summarized and the likely hardware and software categories used by those systems are enumerated. The importance of industry partners is highlighted. This paper also identifies an alternate path forward without industry partners to advance the knowledge needed to use artificial intelligence to analyze HBOMs and SBOMs to better understand reactor resiliency.

cybersecurity↗

A CHIL Validation of Machine Learning-Assisted Methods for Real-Time Controls of Solar PV for Grid Services: Preprint

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been proposed; however, these technologies lack comprehensive validation under real-world application scenarios. This paper addresses this gap by designing and developing a controller-hardware-in-the-loop framework to evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. Simulation results indicate the superior performance of an ML-based approach compared to the conventional reference-control grouping-based approach, showcasing its potential to support grid stability and operational efficiency.

closed-loop validation↗

Hardware Fuzzing with An Emulator

Bugs in digital logic have led to some significant security vulnerabilities. Hardware bugs are particularly troublesome since they cannot be easily patched. Additionally, if the bug is in the root of trust, all trust built upon it can be vulnerable. Traditional testing either require a deep knowledge of the system, creative attack vectors and lots of human interaction. This is not scalable as there are very few engineers that can wear the hat of a designer, a verification engineer, and a cybersecurity expert. Hardware fuzzing is a relatively new research area in dynamic hardware testing. It has proven to be an effective method for discovering bugs, unexpected behaviors, and security vulnerabilities in software. While hardware fuzzing is new to the hardware domain, it has a strong track record in software testing. Fuzzing is a testing technique that randomly mutates the input data to uncover bugs or vulnerabilities in the design. It is especially good at finding corner cases that test engineers can not envision. Another advantage over other dynamic testing techniques is that, if done well, deep knowledge of the design is not required. Additionally, fuzzing scales well. If the system is set up correctly, it can run unsupervised for weeks if necessary. In this work, we propose using hardware fuzzing to improve the input vector generation for an information flow tracking tool. To get reasonable throughput of test vectors, an emulator is targeted as the execution platform. Efficient emulator execution has some specific requirements.

42 ENGINEERING↗

Flow dynamics and heat transfer in simplified battery energy storage systems with heated battery modules

Large-scale energy storage systems (ESSs) composed of batteries show promise in addressing current energy challenges, but dissipation of generated heat is important. Here, this paper focuses on buoyant convective flows in simplified ESS battery racks. Natural convection is not generally the primary cooling strategy but can be important in abnormal scenarios where there is module overheat or potentially thermal runaway. We use computational fluid dynamics to investigate the flow dynamics and heat transfer mechanisms in a simplified parameterized rack design. Despite its simplicity, this configuration produces many of the relevant features expected in real ESSs without details of module geometry or hardware, allowing broad conclusions independent of manufacture-specific designs. We start by providing visualizations of the flowfield and measurements of entrainment, heat flux, and pressure. To characterize the dependence on the system parameters, we develop an integral-scale analysis of the average temperature equation to highlight the dominant source terms. We use results from this analysis to derive a steady network model composed of simple algebraic expressions to provide first-order predictions of entrainment through the rack. The network model leads to a linear scaling of the Reynolds number based on convective mass flux with respect to the Grashof number based on the heat source. We deduce empirical relationships that relate the heat exchanged between modules using a surface-averaged Nusselt number as a function of the local Reynolds and Rayleigh numbers. Lastly, we investigate how space between the modules and rack in the spanwise direction creates flow bypass, resulting in different flow pathways.

Battery thermal management↗