Search NASA⌕ Search

SEARCH · Search NASA

Results for “Accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Development of male-sterile lines of Setaria viridis to accelerate C 4 model plant genetics

Setaria viridis is a diploid C 4 grass in the Poaceae family, notable for its rapid life cycle of 6–8 weeks from sowing to seed—much shorter than the 4–5 months required by crops such as Zea mays and Sorghum bicolor . This fast growth makes S. viridis a valuable model for C 4 crop research. Genetic crosses are essential for studying gene function, but manual crossing is labor-intensive and time-consuming. Here, to address this, we developed a male-sterile line by targeting the S. viridis ortholog of Setaria italica NO POLLEN 1 ( SiNP1 ), which encodes a glucose–methanol–choline oxidoreductase required for pollen exine formation. Using Cas9 and TREX2 -mediated genome editing, we generated SiNP1 knockouts in both the S. viridis ME034V and A10.1 backgrounds that were fully male-sterile. Backcrossing T 0 male-sterile plants to ME034V wild-type followed by selfing yielded a stable BC 1 F 2 line homozygous for a 59 bp deletion in the S. viridis NO POLLEN 1 gene, easily genotyped by PCR and maintained by heterozygous siblings. Using this line, we developed a simple and efficient crossing protocol that eliminates the need for emasculation. This method enables a single person to perform up to 100 crosses per day—compared to 15 using traditional methods—and yields 20–32 F 1 hybrid seeds per panicle with 100% genetic purity. We also quantified pollen flow and outcrossing frequencies under greenhouse conditions to develop optimal bagging strategies and prevent unintended pollination. This resource accelerates genetic research in S. viridis , enhancing its utility as a premier C 4 model for mapping and functional genomics.

C4 research↗

Altered morphology and diffusivity of water confined in MXenes: Machine learning–accelerated computations combined with experiments

Nanoconfined water exhibits unique properties compared to bulk water due to limited quantities, frustrated hydrogen bonding, and surface interactions, which are fundamental for energy storage and transport applications. We integrate machine learning–accelerated ab initio molecular dynamics with x-ray diffraction (XRD) and inelastic neutron scattering (INS) to systematically analyze the thermodynamic and dynamic behavior of water confined between functionalized (-F, -O, and -OH) two-dimensional (2D) Ti 3 C 2 T x MXene layers. As water intercalates between layers, the interlayer spacing exhibits layer-dependent staging characteristics. The water polarization can be flipped by the count and morphology of intercalated molecules interacting with MXene surface groups, resulting in varying electrostatic potential profiles. On the basis of interfacial electrostatic potential, hydrogen bond lifetime, and molecular orientation, we establish a linear combination of exponential model describing water diffusivity. These computational insights align well with experimental x-ray and neutron measurements, suggesting strategies for tuning water morphology and transport by tailoring MXene surface chemistry and water content for electrochemical energy storage and nanofluidic applications.

Tang, Jiawei [Southeast Univ., Nanjing (China)] (O↗

Universal progression of structure and dynamics in colloidal nanocrystal gels during salt-accelerated aging

Controlling the structure and function of colloidal gels requires a detailed understanding of how the various components govern network formation and aging. In particular, molecular additives like salts are widely used to tune interparticle interactions, yet their influence on gelation pathways in complex systems such as colloidal nanocrystal gels remains inadequately understood. Here, we investigate how noncoordinating salts modulate the evolution of gels formed using chemically linked tin-doped indium oxide nanocrystals. Through combined structural, dynamic, and kinetic analyses, we demonstrate that increasing salt concentration accelerates gelation. When rescaled by salt-dependent characteristic times, the evolution collapses onto universal trajectories, revealing a time-salt superposition principle. The universality extends across length scales, suggesting a consistent salt-dependent mechanism that controls both local structuring and macroscopic network formation. This observed salt modulation of structure and dynamics provides a predictive basis for controlling the kinetics of nonequilibrium nanocrystal gel assembly, enhancing the rational design of functional nanomaterials with tunable properties.

36 MATERIALS SCIENCE↗

Genomic approaches to accelerate American chestnut restoration

More than a century after two introduced pathogens killed billions of American chestnut trees, introgression of resistance alleles from Chinese chestnuts has contributed to the recovery of self-sustaining populations. However, progress has been slow because of the complex genetic architecture of resistance. To better understand blight resistance, we compared reference genomes, gene expression responses, and stem metabolite profiles of the resistant Chinese and susceptible American chestnut species. To accelerate resistance breeding, we conducted large-scale phenotyping and genotyping in hybrids of these species. Simulation and inoculation experiments suggest that significant resistance gains are possible through selectively breeding trees with an average of 70 to 85% American chestnut ancestry. In conclusion, the resources developed in this work are foundational for breeding to create diverse restoration populations with sufficient disease resistance and competitive growth.

Westbrook, Jared W. [The American Chestnut Foundat↗

Accelerated Constrained Sparse Tensor Factorization on Massively Parallel Architectures

This study presents the first constrained sparse tensor factorization (cSTF) framework that optimizes and fully offloads computation to massively parallel GPU architectures, and the first performance characterization of cSTF on GPU architectures. In contrast to prior work on tensor factorization, where the matricized tensor times Khatri-Rao product (MTTKRP) is the primary performance bottleneck, our systematic analysis of the cSTF algorithm on GPUs reveals that adding constraints creates an additional bottleneck in the update operation for many real-world sparse tensors. While executing the update operation on the GPU brings significant speedup over its CPU counterpart, it remains a significant bottleneck. To further accelerate the update operation, we propose cuADMM, a new update algorithm that leverages algorithmic and code optimization strategies to minimize both computation and data movement on GPUs. As a result, our framework delivers significantly improved performance compared to prior state-of-the-art. On 10 real-world sparse tensors, our framework achieves geometric mean speedup of 5.1 × (max 41.59 ×) and 7.01 × (max 58.05 ×) on the NIVIDA A100 and H100 GPUs, respectively, over the state-of-the-art SPLATT library running on a 26-core Intel Ice Lake Xeon CPU.

Soh, Yongseok↗

Unveiling the Thermal Stability of Sodium Ion Pouch Cells Using Accelerating Rate Calorimetry

The thermal stability of ~420 mAh Na 0.97 Ca 0.03 [Mn 0.39 Fe 0.31 Ni 0.22 Zn 0.08 ]O 2 (NCMFNZO)/hard carbon (HC) pouch cells was investigated using accelerating rate calorimetry (ARC) at elevated temperatures. 1 m NaPF 6 in propylene carbonate (PC):ethyl methyl carbonate (EMC) (1:1 by volume) was used as a control electrolyte. Adding 2 wt% fluoroethylene carbonate to the electrolyte improves the cell's thermal stability by decreasing the self-heating rate (SHR) across the whole testing temperature range. The selected states-of-charge (SoC), including 70%, 84%, and 100%, exhibit minimal impact on the exothermic behavior, except for a slight decrease in SHR after ~275 °C at 70% SoC. When compared to traditional lithium-ion batteries operating at 100% SoC, NCMFNZO/HC pouch cells demonstrate inferior thermal stability compared to LiFePO 4 (LFP)/graphite pouch cells, displaying a higher SHR from 220 to 300 °C. LiNi 0.8 Mn 0.1 Co 0.1 O 2 /graphite + SiO x pouch cells exhibit the worst safety performance, with an early onset temperature of ~100 °C and the highest SHR across the entire temperature range. These results offer a direct comparison of the impact of SoC and electrolyte compositions on the thermal stability of SIBs at elevated temperatures, highlighting that there is still room for improvement in SIBs safety performance compared to LFP/graphite chemistry.

42 ENGINEERING↗

Miniaturize the Redox Flow Battery for Accelerated Materials Discovery and Development

Redox flow batteries are a promising technology for grid-scale energy storage. The aqueous organic redox flow battery is of particular interest for its potentially low material cost and sustainability. Developing novel organic active material for flow battery electrolytes typically entails molecular engineering toward desired properties, necessitating organic synthesis. In a research laboratory setting, the synthesis of specifically designed organic molecules featuring targeted functional groups is time and resources intensive. In the past, synthesizing materials required for battery testing has often required gram-scale production, presenting considerable constraints on the pace of novel organic material discovery. In this report, we introduce a miniaturized cell design that mandates only milligram-scale material synthesis while yielding testing outcomes equivalent or superior to those reported with other commercially available or homemade flow cells in the literature. The test results under various pH conditions validate the scale-down strategy to accelerate the flow battery material discovery and development using the newly designed mini cell. This approach offers researchers an efficient means to notably reduce the time and resources required to develop novel materials for flow batteries.

25 ENERGY STORAGE↗

Generic Multi-Layer Perceptron Inference Accelerator on FPGA (vneuron) v1.0

We have designed and implemented a neural network inference compute engine (vneuron) that can be deployed in the fabric of any FPGA without using special hardware accelerator primitive. The "vneuron" is purely written in verilog, and supports scalable neural network structure with fully connected layers and ReLU activation ( Multi-Layer Perceptron architecture) with 16 bits of precision. We have demonstrated it on an Xilinx Artix 7 FPGA for a 16-input, 8-output MLP with 3 layer, 1600 parameters. It takes 40 DSP48E and 40 BRAM18, and takes 131 clock cycles for computing (1048 ns when clocked at 125MHz). We include PyTorch quantization from a given floating point model, and provide behavioral verification simulation in the disclosed software package.

Du, Qiang↗

ESnet-JLab FPGA Accelerated Transport (control plane) [EJFAT (udplbd2)] v2.0

The ESnet-JLab FPGA Accelerated Transport system is a solution for streaming high-speed scientific measurement data from Data Acquisition Systems (DAQs) to high-performance computing facilties. It is generally compatible with many science workflows, and makes no assumptions about the specifics of any particular experiment. This program (udplbd version 2) implements the control plane for the system. It is responsible for programming network forwarding rules into the data plane (implemented by the hardware designed named udplb, described separately). It also implements the control loop necessary to match up offered workload with available capacity on high-performance compute nodes.

Howard, Derek [Lawrence Berkeley National Laborato↗

ESnet-JLab FPGA Accelerated Transport (data plane) [EJFAT (udplb)] v1.0

The ESnet-JLab FPGA Accelerated Transport system is a solution for streaming high-speed scientific measurement data from Data Acquisition Systems (DAQs) to high-performance computing facilties. It is generally compatible with many science workflows, and makes no assumptions about the specifics of any particular experiment. This program (udplb) implements the data plane portion of the EJFAT system. It is an FPGA design that rewrites and forwards data packets from a UDP-based scientific workflow to high-performance compute nodes. It depends on another program (udplbd, disclosed separately) to implement the control system.

Bengough, Peter [Malleable Networks, Inc.]↗

Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”

This page contains the datasets and code #O5097 Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”. Data literature-derived cultivation experiments for polyhydroxybutyrate (PHB) production in Synechocystis sp. PCC 6803 and were used for ML model development, interpretation, and experimental validation.

Lalonde, Jessica N. [Los Alamos National Laborator↗

Ginkgo - A math library designed to accelerate Exascale Computing Project science applications

Large-scale simulations require efficient computation across the entire computing hierarchy. A challenge of the Exascale Computing Project (ECP) was to reconcile highly heterogeneous hardware with the myriad of applications that were required to run on these supercomputers. Mathematical software forms the backbone of almost all scientific applications, providing efficient abstractions and operations that are crucial to harness the performance of computing systems. Ginkgo is one such mathematical software library, nurtured by ECP, providing high-performance, user-friendly, and performance portable interfaces for applications in ECP and beyond. In this paper, we elaborate on Ginkgo’s philosophy of high-performance software that is sustainable, reproducible, and easy to use. We showcase the wide feature set of solvers and preconditioners available in Ginkgo and the central concepts involved in their design. We elaborate on four different ECP software integrations: MFEM, PeleLM + SUNDIALS, XGC, and ExaSGD that use Ginkgo to accelerate their science runs. Performance studies of different problems from these applications highlight the effectiveness of Ginkgo and the benefits incurred by these ECP applications.

Cojean, Terry↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

A graphics processing unit accelerated sparse direct solver and preconditioner with block low rank compression

We present the GPU implementation efforts and challenges of the sparse solver package STRUMPACK. The code is made publicly available on github with a permissive BSD license. STRUMPACK implements an approximate multifrontal solver, a sparse LU factorization which makes use of compression methods to accelerate time to solution and reduce memory usage. Multiple compression schemes based on rank-structured and hierarchical matrix approximations are supported, including hierarchically semi-separable, hierarchically off-diagonal butterfly, and block low rank. Here, in this paper, we present the GPU implementation of the block low rank (BLR) compression method within a multifrontal solver. Our GPU implementation relies on highly optimized vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs, rocBLAS and rocSOLVER for AMD GPUs and the Intel oneAPI Math Kernel Library (oneMKL) for Intel GPUs. Additionally, we rely on external open source libraries such as SLATE (Software for Linear Algebra Targeting Exascale), MAGMA (Matrix Algebra on GPU and Multi-core Architectures), and KBLAS (KAUST BLAS). SLATE is used as a GPU-capable ScaLAPACK replacement. From MAGMA we use variable sized batched dense linear algebra operations such as GEMM, TRSM and LU with partial pivoting. KBLAS provides efficient (batched) low rank matrix compression for NVIDIA GPUs using an adaptive randomized sampling scheme. The resulting sparse solver and preconditioner runs on NVIDIA, AMD and Intel GPUs. Interfaces are available from PETSc, Trilinos and MFEM, or the solver can be used directly in user code. We report results for a range of benchmark applications, using the Perlmutter system from NERSC, Frontier from ORNL, and Aurora from ALCF. For a high frequency wave equation on a regular mesh, using 32 Perlmutter compute nodes, the factorization phase of the exact GPU solver is about 6.5× faster compared to the CPU-only solver. The BLR-enabled GPU solver is about 13.8× faster than the CPU exact solver. For a collection of SuiteSparse matrices, the STRUMPACK exact factorization on a single GPU is on average 1.9× faster than NVIDIA’s cuDSS solver.

97 MATHEMATICS AND COMPUTING↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

Coarse Mesh Finite Difference Acceleration for Pebble Tracking Transport in Griffin

We implemented a coarse mesh finite difference (CMFD) for accelerating transport calculations with PTT (pebble tracking transport) in the Griffin code. More specifically, extensions for transport update with the consideration of scattering operator and CMFD projection were implemented for PTT. The implementation was verified with a simplified PBR (pebble bed reactor) benchmark problem and significant performance improvements in CPU time was observed.

97 MATHEMATICS AND COMPUTING↗

Longitudinal shaping of plasma waveguides using diffractive axicons for laser wakefield acceleration

New techniques for the optical generation of plasma waveguides—optical fibers for ultra-intense light pulses—have become vital to the advancement of multi-GeV laser wakefield acceleration. Here, we demonstrate the fabrication and characterization of a transmissive 8-level logarithmic diffractive axicon (LDA) for the generation of meter-scale plasma waveguides. These LDAs enable the formation of a Bessel-like beam with controllable start and end locations of the focal line and near-constant intensity on axis. We present measurements of the Bessel-like focal profile produced by the LDA and of the leading end of the plasma column generated by it. One important feature is the formation of a funnel-mouthed plasma channel entrance that can act as waveguide coupler. We also compare the diffraction efficiency of our 8-level LDA to 4-level and binary versions, with measurements comparing well to theory.

Tripathi, N.↗