Search NASA⌕ Search

SEARCH · Search NASA

Results for “accelerator”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Metric Learning to Accelerate Convergence of Operator Splitting Methods

Recent developments in machine learning have led to promising advances in accelerating the solution of constrained optimization problems. Increasing demand for real-time decision-making capabilities in applications such as artificial intelligence and optimal control has led to a variety of proposed strategies for learning to produce fast solutions to optimization problems. For example, recent works have shown that it is possible to accelerate the convergence of optimization algorithms by learning to select their parameters, such as gradient descent stepsizes. This work proposes a new approach, in which the underlying metric spaces of proximal operator splitting algorithms are learned to maximize convergence rate. While prior works in optimization theory have derived optimal metrics in simple cases, no such result exists for many practical problem forms including general Quadratic Programming (QP). This paper shows how differentiable optimization can enable the end-to-end learning of proximal metrics, enhancing the convergence of proximal algorithms for QP problems beyond what is possible based on known theory. Additionally, the results illustrate a strong connection between the learned proximal metrics and active constraints at the optima, leading to an interpretation in which the predicted proximal metrics can be viewed as a form of active set prediction.

King, Ethan [BATTELLE (PACIFIC NW LAB)]↗

ICED: An Integrated CGRA Framework Enabling DFVS-Aware Acceleration

oarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. Existing CGRA mapping approaches extract instruction-level parallelism, exploit loop-pipelining opportunities, guarantee the data dependency, and target high throughput of a given loop. However, the recurrence data-dependency in the DFG and the mismatch between required and available computing/communication resources complicate the mapping, and might lead to significant unbalances in the utilization of the CGRA's tiles. This results in wasted power for tiles with low utilization. Applying dynamic voltage and frequency scaling (DVFS) can potentially solve this challenge and improve energy efficiency by adjusting voltage and frequency of different tiles independently. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non performance-constraining kernels. This paper proposes ICEDTEA -- an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICEDTEA proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICEDTEA is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICEDTEA improves average utilization by 2.3$\times$ and energy-efficiency by 1.32$\times$ over a conventional CGRA. With streaming applications, ICEDTEA improves energy efficiency by 1.12$\times$ over a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput.

Tan, Cheng↗

Assessment of Accelerated Stress Testing Data for Silicon Photovoltaics Using Tensor Decomposition Methods

In this work, we examine the use of high-order tensor decompositions to analyze degradation pathways emerging from accelerated stress testing of silicon photovoltaic (PV) modules. Matrix-based decompositions are powerful tools for studying two-dimensional data arrays and form the foundation of a host of classical data analysis techniques. Tensors are high-order extrapolations of matrices that are able to account for more parameter dimensions, and a variety of tensor decomposition methods have been developed that similarly seek to extend insights from matrix decompositions to higher dimensions. Applying and interpreting tensor decomposition methods to sequences of PV module image data, we seek to uncover and isolate different degradation modes occurring from accelerated stress testing procedures. Further, we consider the contributions of different modes to PV module performance degradations.

data analysis↗

LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators

Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges, requiring efficient hardware acceleration. Benchmarking the performance of LLMs across diverse hardware platforms is crucial to understanding their scalability and throughput characteristics. We introduce LLM-Inference-Bench, a comprehensive benchmarking suite to evaluate the hardware inference performance of LLMs. We thoroughly analyze diverse hardware platforms, including GPUs from Nvidia and AMD and specialized AI accelerators, Intel Habana and SambaNova. Our evaluation includes several LLM inference frameworks and models from LLaMA, Mistral, and Qwen families with 7B and 70B parameters. Our benchmarking results reveal the strengths and limitations of various models, hardware platforms, and inference frameworks. We provide an interactive dashboard to help identify configurations for optimal performance for a given hardware platform.

Chitty-Venkata, Krishna Teja↗

Study of Shock Formation Parameters With Drive Conditions in Magnetically Accelerated Plasma Flows

We present experimental data regarding the formation of high-energy-density shocks in magnetically accelerated plasma flows using pulsed power drivers. We quantify the flow velocity and temperature of the ablated plasma using optical Thomson scattering and gated emission imaging across two different generators. We show that, regardless of the drive parameters, the plasma flows show continuous acceleration over centimeter spatial scales, in line with trends in published simulation work. When stationary targets are placed in these supersonic flows, bow-shock formation is observed at all drive parameters in a range of materials. In the higher density flow generated on the 1-MA COBRA generator at Cornell University, heating of the upstream flow ahead of the shock is observed and quantified, which is not observed at the lower density flow on the 0.2-MA Bertha driver at UC San Diego. Here, when combined with previous work on the XP generator at Cornell, we can show that these three experimental setups allow control of the effect of radiation loss and upstream absorption on the formation of the bow shock.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Rev1 overexpression accelerates N -methyl- N -nitrosourea (MNU)-induced thymic lymphoma by increasing mutagenesis

Rev1 has two important functions in the translesion synthesis pathway, including dCMP transferase activity, and acts as a scaffolding protein for other polymerases involved in translesion synthesis. However, the role of Rev1 in mutagenesis and tumorigenesis in vivo remains unclear. We previously generated Rev1-overexpressing (Rev1-Tg) mice and reported that they exhibited a significantly increased incidence of intestinal adenoma and thymic lymphoma (TL) after N-methyl-N-nitrosourea (MNU) treatment. In this study, we investigated mutagenesis of MNU-induced TL tumorigenesis in wild-type (WT) and Rev1-Tg mice using diverse approaches, including whole-exome sequencing (WES). In Rev1-Tg TLs, the mutation frequency was higher than that in WT TL in most cases. However, no difference in the number of nonsynonymous mutations in the Catalogue of Somatic Mutations in Cancer (COSMIC) genes was observed, and mutations involved in Notch1 and MAPK signaling were similarly detected in both TLs. Mutational signature analysis of WT and Rev1-Tg TLs revealed cosine similarity with COSMIC mutational SBS5 (aging-related) and SBS11 (alkylation-related). Interestingly, the total number of mutations, but not the genotypes of WT and Rev1-Tg, was positively correlated with the relative contribution of SBS5 in individual TLs, suggesting that genetic instability could be accelerated in Rev1-Tg TLs. Finally, we demonstrated that preleukemic cells could be detected earlier in Rev1-Tg mice than in WT mice, following MNU treatment. In conclusion, Rev1 overexpression accelerates mutagenesis and increases the incidence of MNU-induced TL by shortening the latency period, which may be associated with more frequent DNA damage-induced genetic instability.

60 APPLIED LIFE SCIENCES↗

Accelerating iterative ptychography with an integrated neural network

Electron ptychography is a powerful and versatile tool for high-resolution and dose-efficient imaging. Iterative reconstruction algorithms are powerful but also computationally expensive due to their relative complexity and the many hyperparameters that must be optimised. Gradient descent-based iterative ptychography is a popular method, but it may converge slowly when reconstructing low spatial frequencies. Here, in this work, we present a method for accelerating a gradient descent-based iterative reconstruction algorithm by training a neural network (NN) that is applied in the reconstruction loop. The NN works in Fourier space and selectively boosts low spatial frequencies, thus enabling faster convergence in a manner similar to accelerated gradient descent algorithms. We discuss the difficulties that arise when incorporating a NN into an iterative reconstruction algorithm and show how they can be overcome with iterative training. We apply our method to simulated and experimental data of gold nanoparticles on amorphous carbon and show that we can significantly speed up ptychographic reconstruction of the nanoparticles.

4DSTEM↗

Small-scale production of bespoke accelerated aging plutonium alloy

Here, this study demonstrates the 100 g scale manufacture of a plutonium alloy that ages at an accelerated rate. The resulting alloy ages six times faster than typical weapons-grade plutonium due to the addition of 238 Pu. As a major innovation, the process involved using a partial direct oxide reduction technique. This method was achieved by developing a new, complex geometry stirrer using additive manufacturing to reduce the 238Pu oxide and efficiently incorporate it into weapons-grade plutonium metal. The material was then purified by molten salt extraction and electrorefining before being alloyed with gallium. The alloy was then cold-rolled and annealed in a homogenization heat treatment. The resulting disk was characterized by metallography and differential scanning calorimetry, and the impurity content was determined using analytical chemistry techniques. The results show that a homogeneous delta phase plutonium alloy was achieved with expected microstructure and minimal impurities. This study was also successful in changing the plutonium isotopic composition by incorporating additional 238 Pu to accelerate the effects of radiation damage. This enables researchers to study long-term aging phenomena in a reduced time frame, thus avoiding the need for large-scale material production and circumventing the limitations of using naturally aged, archived plutonium.

Aging theory↗

Mechanochemically accelerated deconstruction of chemically recyclable plastics

Plastics redesign for circularity has primarily focused on monomer chemistries enabling faster deconstruction rates concomitant with high monomer yields. Yet, during deconstruction, polymer chains interact with their reaction medium, which remains underexplored in polymer reactivity. Here, we show that, when plastics are deconstructed in reaction media that promote swelling, initial rates are accelerated by over sixfold beyond those in small-molecule analogs. This unexpected acceleration is primarily tied to mechanochemical activation of strained polymer chains; however, changes in the activity of water under polymer confinement and bond activation in solvent-separated ion pairs are also important. Together, deconstruction times can be shortened by seven times by codesigning plastics and their deconstruction processes.

36 MATERIALS SCIENCE↗

Two-Level Sketching Alternating Anderson Acceleration for Complex Physics Applications

We present a novel two-level sketching extension of the Alternating Anderson–Picard (AAP) method for accelerating fixed-point iterations in challenging single- and multiphysics simulations governed by discretized PDEs. Our approach combines a static, physics-based projection that reduces the least-squares (LS) problem to the most informative field (e.g., via Schur-complement insight) with a dynamic, algebraic sketching stage driven by a backward stability analysis under Lipschitz continuity. We introduce inexpensive estimators for stability thresholds and cache-aware randomized selection strategies to balance computational cost against memory access overhead. The resulting algorithm solves reduced LS systems in place, minimizes memory footprints, and seamlessly alternates between low-cost Picard updates and Anderson mixing. Implemented in Julia, our two-level sketching AAP achieves up to 50% time-to-solution reductions compared to standard Anderson acceleration—without degrading convergence rates—on benchmark problems including Stokes, 𝑝-Laplacian, bidomain, and Navier–Stokes formulations at varying problem sizes. These results demonstrate the method’s robustness, scalability, and potential for integration into high-performance scientific computing frameworks. Our implementation is available open source in the AAP.jl library.

Barnafi, Nicolas [University of Chile, Santiago]↗

Secure API-Driven Research Automation to Accelerate Scientific Discovery

The Secure Scientific Service Mesh (S3M) provides API-driven infrastructure to accelerate scientific discovery through automated research workflows. By integrating near real-time streaming capabilities, intelligent workflow orchestration, and fine-grained authorization within a service mesh architecture, S3M enables secure and flexible programmatic access to high performance computing (HPC) resources. This framework allows intelligent agents and experimental facilities to dynamically provision resources and execute complex workflows, accelerating experimental lifecycles, and enabling AI-augmented autonomous science. S3M establishes a modern foundation for scientific computing infrastructure that significantly reduces traditional barriers between researchers, computational resources, and experimental facilities.

Skluzacek, Tyler [ORNL] (ORCID:0000000322424931)↗

The Memory Scaling of Reverse-Mode Differentiation in Particle Accelerator Simulations with Space Charge

The recent development of differentiable simulation codes for particle accelerators has enabled gradient-based workflows that promise finer control and more realistic modeling of accelerator facilities. However, when using reverse-mode automatic differentiation, the memory usage continuously increases during the simulation, and can potentially exceed the available hardware memory - especially when costly space charge computation is included. To study the memory requirements for differentiable simulations, we have implemented space charge in Cheetah, a PyTorch-based beam tracking code that supports reverse-mode differentiation. We find that the memory usage for reverse-mode differentiation grows linearly with the number of macroparticles and cells, and that it is proportional to the number of space charge kicks involved in the simulation. This general scaling can be used to evaluate whether a given differentiable simulation is feasible given hardware memory constraints.

Dhamrait, Arjun↗

Multiscale Molecular Dynamics Simulations: Accelerating Conformational Sampling of Biomolecular Systems by Iterating All-Atom and Coarse-Grained Simulations

We developed the atomistic-coarse-grained multiscale MD simulation method in the OpenMM simulation package by iterating between the all-atom (AA) and coarse-grained (CG) MD simulations to enhance the sampling of biomolecular conformations. As the free energy surfaces are flattened during CG MD simulations, we can accelerate the transitions between different low-energy conformations. The AA-CG-AA cycles are repeated, facilitating the accelerated sampling of biomolecular conformations at a CG level, while the finer atomistic interactions are refined with AA simulators.

Do, Hung Nguyen↗

Standardizing UI/UX across accelerator labs

During February 26–28, 2025, the first-ever particle accelerator user interface/user experience (UI/UX) workshop was held at SLAC. Attendees had backgrounds ranging from software development to control systems management and human factors (HF) science. The workshop began with participants discussing the current state of UI/UX procedures and practices at their respective laboratories to share experiences and learn from one another. Additional discussions focused on how to effectively integrate UI/UX best practices into actionable goals for developers, managers, and operators when working on new or existing interfaces. The goal of the working group is to create a website that will guide developers, managers, scientists, and end users at accelerator laboratories in incorporating UI/UX best practices into software development. The working group continues to meet virtually toward this goal, and is planning a second workshop for next year.

Tran, Tiffany [SLAC]↗

An Agile/XP software development process for modernizing the accelerator control system at Fermilab

Fermilab is undergoing the most ambitious upgrade to its accelerator control system of the 21st century. As part of the ACORN project, hundreds of legacy control system applications written in C/C++ will be re-imagined and developed from the ground up. In addition, applications to support Fermilab’s new super-conducting linear accelerator are already under construction. To manage the development of modern controls applications, the Controls department has adopted an Agile software development process based on eXtreme Programming. In this paper we will describe our process and detail our experience applying it to the development of two case studies.

Diamond, John [Fermilab]↗

Accelerator Complex Evolution at Fermilab

The largest hadron accelerator facility in the US is undergoing radical changes and the undertaking of new HEP-driven neutrino research. This talk will discuss the wide-ranging projects and impacts to the accelerator community taking place at FNAL.

Convery, M. [Fermilab]↗

First Results from a Nb3Sn-Coated 1.5-Cell 650 MHz SRF Cavity for Cryogen-Free Industrial Accelerators

First Results from a Nb3Sn-Coated 1.5-Cell 650 MHz SRF Cavity for Cryogen-Free Industrial Accelerators ABSTRACT = Fermilab is advancing the development of a compact, high-power electron beam accelerator using superconducting radio frequency (SRF) technology as a non-radioactive alternative to traditional radiological sources. The current design targets continuous-wave (CW) operation at \SI{1.6}{MeV} and \SI{20}{kW}. To ensure suitability for industrial environments, the system is being designed for cryogen-free operation, driving the adoption of a novel Nb$_3$Sn-coated 1.5-cell SRF cavity operating at \SI{650}{MHz}. This contribution reports on the fabrication, surface preparation, and Nb$_3$Sn coating process of the cavity, as well as first results from vertical test stand (VTS) measurements performed in a liquid helium bath. These initial tests mark a key milestone toward demonstrating the viability of conduction-cooled Nb$_3$Sn SRF cavities for industrial-scale deployment.

Tagdulang, N. [Fermilab]↗

High-Fidelity Accelerated Design of High-performance Electrochemical Systems

Large-scale electrification is vital to addressing the climate crisis, but several scientific and technological challenges remain to fully electrify both the chemical industry and transportation. In both of these areas, new electrochemical materials will be critical, but their development currently relies heavily on human-time-intensive experimental trial and error and computationally expensive first-principles, meso-scale and continuum simulations. To accelerate this process, our team has developed the AutoMat platform. AutoMat can accelerate development of new electrochemical materials along two avenues: first, automated input generation and management of simulations at multiple lengthscales as well as “handoff” of outputs from one lengthscale as inputs to the next; and second, replacement of the most computationally intensive simulation processes with machine-learned surrogate models. The crux of our team’s effort was not “reinventing the wheel” by developing entirely new techniques, but rather building a “superhighway” that allows existing state-of-the-art techniques to run faster and more smoothly than before. AutoMat can utilize tools spanning from first-principles quantum chemistry computations to automated robotic experimentation, and is driven by design space search techniques to reduce the number of iterations through the full simulation loop by rapidly targeting promising regions of design spaces such as single-atom alloy catalysts or blends of liquid electrolytes.

25 ENERGY STORAGE↗