Search NASA⌕ Search

SEARCH · Search NASA

Results for “Accelerator Design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Modeling and simulation of multiphase flows

This presentation provides an overview of the National Energy Technology Laboratory’s (NETL) multiphase computational fluid dynamics codes. The highly successful Multiphase Flows with Interphase eXchanges (MFIX) suite has been used to model a wide range of applications including post-combustion carbon capture, bioreactor optimization, and bio-FCC regeneration. MFIX-Exa, a state-of-the-art CFD code, developed under DOE’s Exascale Computing Project, is built on the AMReX software framework (https://amrex-codes.github.io/) and is designed to leverage modern accelerator-based compute architectures. This presentation further reviews the underlying physical models of both MFIX and MFIX-Exa and contrasts their similarities and differences. Examples of past and present CFD simulations will illustrate how scientific computing at NETL is being used not only for scientific exploration but also for design, optimization and scale-up of multiphase flow devices.

Musser, Jordan [NETL]↗

Performance Impact and Trade-Offs for Tuning Key Architectural Parameters on CPU+GPU Systems

In this work, we performed an initial design space exploration of an accelerated processing unit (APU)—a hybrid CPU+GPU architecture that integrates both compute units (CUs) and memory into a unified system. This integration aims to reduce data movement, enhance memory locality, and improve energy efficiency by enabling the CPU and GPU to share memory directly. This effort focused on the interplay of key design components—cache line size, the number of CUs, and main memory technology—and the trade-offs of each configuration were analyzed. This paper highlights the various configurations’ impact on memory accesses, data reuse, and power utilization. The results provide valuable insights that can be leveraged to optimize APU architectures for high-performance and energy-efficient computing and thus create a balanced architecture. This optimization can be achieved by adopting dynamic cache management, runtime CU scaling, and advanced memory integration, highlighting the potential of APUs to address critical challenges in compute, data movement, and memory power consumption.

Asifuzzaman, Kazi [ORNL] (ORCID:0000000240044791)↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Organic Direct-Bonded-Copper-Based Rapid Prototyping for Silicon Carbide Power Module Packaging

Silicon carbide (SiC) power devices are playing ever- growing roles in high-power-density power electronics converters by offering benefits such as high voltage rating, fast transients, and high thermal performance. Organic direct-bonded copper (ODBC)-based packaging, due to its ductility and ease of han- dling, allows the possibility of a more flexible layout design that may better tap the potential of SiC benefits. In this work, an ODBC-based prototyping routine is developed that accelerates the iterations of packaging layout design with low cost. The properties of ODBC and its handling are briefly introduced, and tools and fabrication steps are explained. Following this routine, a 1.2-kV SiC half-bridge power module is designed and fabricated with the focus on sub-nanohenry ultra-low loop inductance. Simulation and experimental validation are also conducted.

25 ENERGY STORAGE↗

HEP High Power Targetry Roadmap -- Workshop Report

Designing a reliable target is already a challenge for MW-class facilities today and has led several major accelerator facilities to operate at lower than design power due to target concerns. With present plans to increase beam power for next generation accelerator facilities in the next decade, timely R and D in support of robust high power targets is critical to secure the full physics benefits of ambitious accelerator power upgrades. A comprehensive R and D program must be implemented to address the many complex challenges faced by multi MW beam intercepting devices. This roadmap is envisioned to be helpful to the DOE-OHEP office when planning and prioritizing future R and D activities as well as leveraging synergies across the Office of Science. The roadmap will be extremely beneficial to the broader (external to DOE HEP) HPT community by communicating OHEP s high level strategy and objectives for HPT R and D and highlighting possible opportunities for collaboration.

43 PARTICLE ACCELERATORS↗

Computational study of tungsten and depleted uranium photoneutron targets for a 20 MeV electron linear accelerator

Neutron production can be realized with a high energy electron linear accelerator by using Bremsstrahlung and photoneutron converters. In this study, Monte Carlo N-Particle Code (MCNP) was used to evaluate potential photonuclear target designs for a high energy electron linear accelerator for applications such as neutron radiography and neutron resonance spectroscopy. A computational model was developed to inform a target design that would yield a high number of neutrons. It consists of a 20 MeV electron beam incident on a Bremsstrahlung target and a photonuclear target to generate neutrons. This computational model showed that a thickness of 0.75 inches for both tungsten and depleted uranium yields the most neutrons from photoneutron reactions. Saturation in the total number of generated neutrons was observed at over 0.75-inch thickness for both evaluated materials. Depleted uranium yielded approximately twice the number of neutrons overall compared to tungsten. The highest neutron surface flux for Depleted Uranium was 1.06 × 10-4 neutrons/cm2/source electron, and for Tungsten it was 5.12 × 10-5 neutrons/cm2/source electron. The optimal target design for this study’s application would consist of a 0.75 inch-thick block of depleted uranium with the length, width, and/or diameter varying dependent on application.

43 PARTICLE ACCELERATORS↗

Machine learning for accuracy in density functional approximations

Machine learning techniques have found their way into computational chemistry as indispensable tools to accelerate atomistic simulations and materials design. In addition, machine learning approaches hold the potential to boost the predictive power of computationally efficient electronic structure methods, such as density functional theory, to chemical accuracy and to correct for fundamental errors in density functional approaches. In this paper, recent progress in applying machine learning to improve the accuracy of density functional and related approximations is reviewed. Promises and challenges in devising machine learning models transferable between different chemistries and materials classes are discussed with the help of examples applying promising models to systems far outside their training sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Milestone in predicting core plasma turbulence: successful multi-channel validation of the gyrokinetic code GENE

On the basis of several recent breakthroughs in fusion research, many activities have been launched around the world to develop fusion power plants on the fastest possible time scale. In this context, high-fidelity simulations of the plasma behavior on large supercomputers provide one of the main pathways to accelerating progress by guiding crucial design decisions. When it comes to determining the energy confinement time of a magnetic confinement fusion device, which is a key quantity of interest, gyrokinetic turbulence simulations are considered the approach of choice – but the question, whether they are really able to reliably predict the plasma behavior is still open. The present study addresses this important issue by means of careful comparisons between state-of-the-art gyrokinetic turbulence simulations with the GENE code and experimental observations in the ASDEX Upgrade tokamak for an unprecedented number of simultaneous plasma observables.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Active learning of ternary alloy structures and energies

Abstract Machine learning models with uncertainty quantification have recently emerged as attractive tools to accelerate the navigation of catalyst design spaces in a data-efficient manner. Here, we combine active learning with a dropout graph convolutional network (dGCN) as a surrogate model to explore the complex materials space of high-entropy alloys (HEAs). We train the dGCN on the formation energies of disordered binary alloy structures in the Pd-Pt-Sn ternary alloy system and improve predictions on ternary structures by performing reduced optimization of the formation free energy, the target property that determines HEA stability, over ensembles of ternary structures constructed based on two coordinate systems: (a) a physics-informed ternary composition space, and (b) data-driven coordinates discovered by the Diffusion Maps manifold learning scheme. Both reduced optimization techniques improve predictions of the formation free energy in the ternary alloy space with a significantly reduced number of DFT calculations compared to a high-fidelity model. The physics-based scheme converges to the target property in a manner akin to a depth-first strategy, whereas the data-driven scheme appears more akin to a breadth-first approach. Both sampling schemes, coupled with our acquisition function, successfully exploit a database of DFT-calculated binary alloy structures and energies, augmented with a relatively small number of ternary alloy calculations, to identify stable ternary HEA compositions and structures. This generalized framework can be extended to incorporate more complex bulk and surface structural motifs, and the results demonstrate that significant dimensionality reduction is possible in thermodynamic sampling problems when suitable active learning schemes are employed.

Chemistry↗

Using the Metropolis algorithm to explore the loss surface of a recurrent neural network

In the limit of small trial moves the Metropolis Monte Carlo algorithm is equivalent to gradient descent on the energy function in the presence of Gaussian white noise. This observation was originally used to demonstrate a correspondence between Metropolis Monte Carlo moves of model molecules and overdamped Langevin dynamics, but it also applies in the context of training a neural network: making small random changes to the weights of a neural network, accepted with the Metropolis probability, with the loss function playing the role of energy, has the same effect as training by explicit gradient descent in the presence of Gaussian white noise. We explore this correspondence in the context of a simple recurrent neural network. We also explore regimes in which this correspondence breaks down, where the gradient of the loss function becomes very large or small. In these regimes the Metropolis algorithm can still effect training, and so can be used as a probe of the loss function of a neural network in regimes in which gradient descent struggles. We also show that training can be accelerated by making purposely-designed Monte Carlo trial moves of neural-network weights.

Casert, Corneel↗

Equilibrium Core Model for Micro Pebble Bed Reactors Using OpenMC

Estimating the equilibrium state for pebble bed reactors (PBRs) presents complex challenges as it requires simultaneous consideration of changes in the pebbles’ movement as well as their fuel compositions. Whereas traditional approaches use multigroup diffusion codes for neutronics calculations of PBRs’ equilibrium state, the double-heterogeneity of PBRs complicates neutron cross-section generation. Continuous-energy Monte Carlo (MC) methods are better suited for detailed PBR analysis because of their natural handling of double-heterogeneity, but they demand substantially more computational resources. Here, this study introduces a novel method for efficiently estimating the equilibrium state in small and micro PBRs with reduced computational cost. The method is anticipated to accelerate the processes of core design and performing parametric studies for utilizing advanced fuel and structural materials. The HTR-10 reactor design was used for validating the method’s predictions and evaluating its computational efficiency. When compared to reference calculation values from the literature, criticality (k-effective) was predicted to be approximately within the margin of error of the MC transport calculation, average core power density (in megawatts per cubic meter) was predicted within 2.5% relative error, and maximum thermal flux (10 13 n/cm 2 .s −1 ) was predicted within 1.8% relative error. The calculated inventory of fission products and fuel composition in the equilibrium core were within 15% and 16.6%, respectively, when compared to reported values from the literature. The difference is attributed to variance in the considered values of the core temperature, which was found to significantly affect the depletion analyses.

Equilibrium core↗

The ATLAS experiment at the CERN Large Hadron Collider: a description of the detector configuration for Run 3

The ATLAS detector is installed in its experimental cavern at Point 1 of the CERN Large Hadron Collider. During Run 2 of the LHC, a luminosity of ℒ = 2 × 10 34 cm -2 s -1 was routinely achieved at the start of fills, twice the design luminosity. For Run 3, accelerator improvements, notably luminosity levelling, allow sustained running at an instantaneous luminosity of ℒ = 2 × 10 34 cm -2 s -1 , with an average of up to 60 interactions per bunch crossing. The ATLAS detector has been upgraded to recover Run 1 single-lepton trigger thresholds while operating comfortably under Run 3 sustained pileup conditions. A fourth pixel layer 3.3 cm from the beam axis was added before Run 2 to improve vertex reconstruction and b-tagging performance. New Liquid Argon Calorimeter digital trigger electronics, with corresponding upgrades to the Trigger and Data Acquisition system, take advantage of a factor of 10 finer granularity to improve triggering on electrons, photons, taus, and hadronic signatures through increased pileup rejection. The inner muon endcap wheels were replaced by New Small Wheels with Micromegas and small-strip Thin Gap Chamber detectors, providing both precision tracking and Level-1 Muon trigger functionality. Trigger coverage of the inner barrel muon layer near one endcap region was augmented with modules integrating new thin-gap resistive plate chambers and smaller-diameter drift-tube chambers. Tile Calorimeter scintillation counters were added to improve electron energy resolution and background rejection. Upgrades to Minimum Bias Trigger Scintillators and Forward Detectors improve luminosity monitoring and enable total proton-proton cross section, diffractive physics, and heavy ion measurements. These upgrades are all compatible with operation in the much harsher environment anticipated after the High-Luminosity upgrade of the LHC and are the first steps towards preparing ATLAS for the High-Luminosity upgrade of the LHC. This paper describes the Run 3 configuration of the ATLAS detector.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multitarget Rydberg gates via spatial blockade engineering

Multi-target gates offer the potential to reduce gate depth in syndrome extraction for quantum error correction. Although neutral-atom quantum computers have demonstrated native multi-qubit gates, existing approaches that avoid additional control or multiple atomic species have been limited to single-target gates. We propose single-control-multi-target CZ^n gates on a single-species neutral-atom platform that require no extra control and have gate durations comparable to standard CZ gates. Our approach leverages tailored interatomic distances to create an asymmetric blockade between the control and target atoms. Using a GPU-accelerated pulse synthesis protocol, we design smooth control pulses for CZZ and CZZZ gates, achieving fidelities of up to 99.55% and $99.24\%$, respectively, even in the presence of simulated atom placement errors and Rydberg-state decay. Our approach is most effective for N=2 (CZZ) and N=3 targets (CZZZ); for larger N, increasing spatial crowding of the targets introduces significant challenges for maintaining the required blockade asymmetry. This work presents a practical path to implementing low-overhead multi-target gates in single-species neutral-atom systems, significantly reducing the resource overhead for syndrome extraction. To motivate the impact of these gates, we apply a greedy scheduling algorithm and we demonstrate that our proposed gates can reduce the number of atom reconfiguration costs by up to 50% for color code syndrome extraction of code distances greater than 5.

Stein, Samuel A.↗

Spectroelectrochemical insights into the intrinsic nature of lead halide perovskites

Abstract Lead halide perovskites have emerged as a new class of semiconductor materials with exceptional optoelectronic properties, sparking significant research interest in photovoltaics and light-emitting diodes. However, achieving long-term operational stability remains a critical hurdle. The soft, ionic nature of the halide perovskite lattice renders them vulnerable to various instabilities. These instabilities can be triggered by factors such as photoexcitation, electrical bias, and the surrounding electrolyte/solvent or atmosphere under operating conditions. Spectroelectrochemistry offers a powerful approach to bridge the gap between electrochemistry and photochemistry (or spectroscopy), by providing a comprehensive understanding of the band structure and excited-state dynamics of halide perovskites. This review summarizes recent advances that highlight the fundamental principles, the electronic band structure of halide perovskite materials, and the photoelectrochemical phenomena observed upon photo- and electro-chemical charge injections. Further, we discuss halide instability, encompassing halide oxidation, vacancy formation, ion migration, degradation, and sequential expulsion under electrical bias. Spectroelectrochemical studies that provide a deeper understanding of interfacial processes and halide mobility can pave the way for the design of more robust perovskites, accelerating future research and development efforts. Graphical Abstract

Min, Seonhong↗

Experimental measurements for extracting nonlinear invariants

Nonlinear integrable optics are a promising alternative approach to lattice design. The integrable optics test accelerator (IOTA) at Fermilab has been constructed for dedicated studies of magnetostatic elliptical elements as described by Danilov and Nagaitsev. The most compelling verification of correct implementation of the NIO lattice is direct observation of the analytically expected invariants. This report outlines the experimental and analytical methods for extracting the nonlinear invariants of motion from data gathered in the last IOTA run.

43 PARTICLE ACCELERATORS↗

Enhancements to Lab-on-a-Fish Technology: Final Report for I3T Project 81245

PNNL has engaged in discussions with several organizations regarding potential future collaborations. PNNL has developed the Lab-on-a-Fish prototypes and produced video instructions for an organization conducting research on shark eggs. Several organizations have expressed interest in using the Lab-on-a-Fish for their studies. We’ve also updated two versions of the Lab-on-a-Fish to expand its applications to a wider range of animal studies and accelerate its commercialization. The new design of the Lab-on-a-Fish, which incorporates PCB-based electrodes, has demonstrated the ability to capture clear ECG signals without the need for additional needle-shaped electrodes. This innovation eliminates the need for implanting electrodes beneath the fish's skin at specific locations, which is required with the current version of Lab-on-a-Fish. This change significantly simplifies the manufacturing and implantation process. For the Lab-on-a-Fish design that uses an optical pulse oximeter, the measurement results were primarily influenced by respiratory activity rather than heart rate. This occurred because the oximeter was placed beneath the operculum. Further research and development are needed to explore more suitable locations and methods for accurate pulse oximeter measurements.

59 BASIC BIOLOGICAL SCIENCES↗

PSEC5 Physical Verification

Ultra-Fast timing detectors, having a time resolution of 10pSec or less, are becoming increasingly important for many scientific disciplines, including High Energy Physics, Nuclear Physics, Medical Imaging and more. When developing ultra-fast detectors such as MCPs, LAPPDs, LGADs, SNSPDs, it is necessary to have access to a readout system that is low power, can support many channels and is cost effective. The PSEC5 chip was an application specific integrated circuit (ASIC) developed by the pioneering group of Prof. Henry Frisch at University of Chicago and micro electronic engineers at Fermi National Accelerator Laboratory. Three boards were designed, two to implement the readout system for an LAPPD and one to only characterize the PSEC1 chip. The main goal of the project shifted to setting up a testing environment to focus on the characterization of the chip. My student and I tested and characterized this ASIC using the third chip-on-board printed circuit board. Results show the successful fabrication and operation of some of the key building blocks and lead to the suggestions on modification for the second iteration of the chip.

Rico Aniles, H. D. [Unlisted]↗

Uncovering Structure–Conductivity Relationships in Anion Exchange Membranes (AEMs) Using Interpretable Machine Learning

Anion exchange membranes (AEMs) play a vital role in the performance of water electrolyzers and fuel cells, yet their discovery and optimization remain challenging due to the complexity of structure–property relationships. In this study, we introduce a machine learning framework that leverages conditional graph neural networks (cGNNs) and descriptor-based models and a hybrid graph neural network (HGARE) to predict and interpret ionic conductivity. The descriptor-based pipeline employs principal component analysis (PCA), ablation, and SHAP analysis to identify factors governing anion conductivity, revealing electronic, topological, and compositional descriptors as key contributors. Beyond prediction, dimensionality reduction and clustering are performed by employing t-SNE and KMeans as well as SOM, which reveal distinct membranes clusters, some of which were enriched with high anion conductivity. Among graph-based approaches, the graph convolutional (GCN) achieved strong predictive performance, while the Hybrid Graph Autoencoder-Regressor Ensemble (HGARE) achieved the highest accuracy. Additionally, atom-level saliency maps from GCN provide spatial explanations for conductive behavior, revealing the importance of polarizable and flexible regions. This work contributes to the accelerated and data-driven design of high-performance AEMs.

Naghshnejad, Pegah [Department of Chemical Enginee↗