Search NASA⌕ Search

SEARCH · Search NASA

Results for “graphics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Accelerating high-order continuum kinetic plasma simulations using multiple GPUs

Kinetic plasma simulations solve the Vlasov-Poisson or Vlasov-Maxwell equations to evolve scalar-variable distribution functions in position-velocity phase space and vector-variable electromagnetic fields in configuration space. The immense computational cost of evolving high-dimensional variables, and their large number of degrees of freedom, often limits the utility of continuum kinetic simulations and presents a challenge when it comes to accurately simulating real-world physical phenomena. To address this challenge, we present techniques that accelerate and minimize the computational work required for a scalable Vlasov-Poisson solver. We show theoretical hardware compute and communication bounds for solving a fourth-order finite-volume Vlasov-Poisson system. These bounds are then used to inform and evaluate the design of performance portable algorithms for a multiple graphics processing unit (GPU) accelerated version of the Vlasov-Poisson solver VCK-CPU [1]. We demonstrate that the multi-GPU Vlasov solver implementation, VCK-GPU, simultaneously minimizes required inter-process data transfer while also being bounded by the machine network performance limits. This results in an overall strong scaling speedup per timestep of up to 40x in three-dimensional phase space (one position, two velocity coordinates) and 54x in four dimensional phase space (two position, two velocity coordinates) and a 341x increase in simulation throughput of the GPU accelerated code over the existing CPU code. The GPU code is also able to weak scale up to 256 compute nodes and 1024 GPUs. In conclusion, we demonstrate that the improved compute performance enables exploring configurations which were previously computationally infeasible, including resolving fine-scale distribution function filamentation and multi-species dynamics with realistic electron-proton mass ratios.

Continuum kinetics↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗

Constellation: The autonomous control and data acquisition system for dynamic experimental setups

The operation of instruments and detectors in laboratory or beamline environments presents a complex challenge, requiring stable operation of multiple concurrent devices, often controlled by separate hardware and software solutions. These environments frequently undergo modifications, such as the inclusion of different auxiliary devices depending on the experiment or facility, adding further complexity. The successful management of such dynamic configurations demands a flexible and robust system capable of controlling data acquisition, monitoring experimental setups, enabling seamless reconfiguration, and integrating new devices with limited effort. This paper presents Constellation, a flexible and network-distributed control and data acquisition software framework tailored to laboratory and beamline environments, that addresses the limitations of existing solutions. The framework is designed with a focus on extensibility, providing a streamlined interface for instrument integration. It supports efficient system setup via network discovery mechanisms, promotes stability through autonomous operational features, and provides comprehensive documentation and supporting tools for operators and application developers such as controllers and logging interfaces. At the core of the architectural design is the autonomy of the individual components, called satellites, which can make independent decisions about their operation and communicate these decisions to other components. This paper introduces the design principles and framework architecture of Constellation, presents the available graphical user interfaces, shares insights from initial successful deployments, and provides an outlook on future developments and applications.

Autonomy↗

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

A High-Performance Discrete-Element Framework for Simulating Flow and Jamming of Moisture Bearing Biomass Feedstocks

We developed and verified a high-performance open-source discrete element method (DEM) solver with simultaneously-supported feedstock-specific interaction models, including bonded-sphere, liquid bridge, cohesion, and non-linear contact models. Our solver uses parallel data structures on hybrid central and graphics processing unit (CPU/GPU) architectures, with favorable strong scaling performance observed for large problem sizes comprised of (100 M particles), and 4X single-node GPU speedup. The particles for corn stover feedstock were conceptualized and calibrated based on experimental measurements and results. Sensitivity analyses demonstrate that the mass flow rate from a wedge hopper is governed primarily by moisture content, friction coefficient, and cohesion energy density. The model is used to reproduce experimentally observed hopper jamming results, highlighting that the experimental no-flow trends can only be achieved by using non-spherical particles, liquid bridge and cohesion models, highlighting the importance of using concurrent feedstock specialized models for the effective representation of biomass material handling problems.

bioenergy↗

A Review of the Lawrence Livermore Nuclear Accident Dosimeter 1980s-present

A Nuclear Accident Dosimetry program is a federal requirement for all facilities that have the potential to have a criticality accident. Personnel Nuclear Accident Dosimeter (PNAD) theory and analytical procedures are driven by various scientific needs and interacting regulations. A brief history of the status of USA Department of Energy (DOE) nuclear accident dosimetry regulations, recommendations, and performance testing criteria are given. Then, the history of the Lawrence Livermore National Laboratory (LLNL) PNAD is explored, including changes in the physical dosimeter and adjustments of the analysis method through the last four decades. Finally, the performance of LLNL’s PNAD at criticality accident intercomparison training exercises since 2009 is explored. In general, reported neutron doses have been within or close to DOE-STD-1098 performance criteria while reported gamma doses have been outside of DOE-STD-1098 performance criteria. Reported total absorbed doses have varied in meeting ANSI/HPS N13.3 and ANSI/HPS N13.3 (R2019) performance criteria. Dosimetry staff retirement and turnover have left historical knowledge gaps, yet provided opportunities within the NAD program at LLNL. This review paper serves as an overview of the history and status of the NAD program. Brief technical, procedural and programmatic recommendations to improve LLNL’s NAD program are given. Technical recommendations include investigating orientation factors through modeling or empirical experimentation, investigating gamma dosimetry methods for high-dose scenarios, and exploring other dosimetric methods for simpler, quicker NAD analysis. Procedural recommendations include better documentation of conversion factor (activity-to-fluence and fluence-to-dose) derivations and spectrum uses, and updated analysis spreadsheets or simple Graphic User Interfaces for dose calculations. In conclusion, programmatic recommendations include formalized training for NAD analysts, and having multiple SMEs trained on the NAD program.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Heat transfer coefficients of moving particle beds from flow-dependent thermal conductivity and near-wall resistance

Accurate determination of heat transfer coefficients for flowing packed particle beds is essential to the design of particle heat exchangers and other thermal and thermochemical equipment. While such dense granular flows mostly fall into the well-known plug-flow regime, the discrete nature of granular materials alters the thermal transport processes in both the near-wall and bulk regions of flowing particle beds from their stationary counterparts. As a result, heat transfer correlations based on the stationary particle bed thermal conductivity could be inadequate for flowing particles in a heat exchanger. Most earlier works have achieved a reasonable agreement with experiments by treating granular heat transfer media as a plug-flow continuum with a near-wall thermal resistance in series. However, the thermal conductivity values of the continuum were often obtained from measurements on stationary beds owing to the difficulty of flowing bed measurements. In this work, it was found that the properties of a stationary bed are highly sensitive to the method of particle packing and there is a decrease in the particle bed thermal conductivity and increase in the near-wall thermal resistance, measured as an effective air gap thickness, on the onset of particle flow. These variations in thermal conductivity of stationary and flowing particle beds can lead to errors in heat transfer coefficient calculations. Therefore, the heat transfer coefficients for granular flows were calculated using experimentally determined flowing particle bed thermal conductivity and near-wall air gap for ceramic particles – CARBO CP 40/100 (mean diameter = 275 µm), HSP 40/70 (404 µm) and HSP 16/30 (956 µm); at velocities of 5–15 mm·s –1 ; and temperatures of 300–650 °C. The thermal conductivity and air gap values for CP 40/100 and HSP 40/70 were further used to calculate heat transfer coefficients across different particle bed temperatures and velocities for different parallel-plate heat exchanger dimensions. Furthermore, these calculations, which show good agreement with measured HTC values reported in literature, can be used as a guide for heat exchanger designs. Graphical abstract

14 SOLAR ENERGY↗

GX: a GPU-native gyrokinetic turbulence code for tokamak and stellarator design

GX is a code designed to solve the nonlinear gyrokinetic system for low-frequency turbulence in magnetized plasmas, particularly tokamaks and stellarators. In GX, our primary motivation and target is a fast gyrokinetic solver that can be used for fusion reactor design and optimization along with wide-ranging physics exploration. Here, this has led to several code and algorithm design decisions, specifically chosen to prioritize time to solution. First, we have used a discretization algorithm that is pseudospectral in the entire phase space, including a Laguerre–Hermite pseudospectral formulation of velocity space, which allows for smooth interpolation between coarse gyrofluid-like resolutions and finer conventional gyrokinetic resolutions and efficient evaluation of a model collision operator. Additionally, we have built GX to natively target graphics processors (GPUs), which are among the fastest computational platforms available today. Finally, we have taken advantage of the reactor-relevant limit of small $\rho _*$ by using the radially local flux-tube approach. In this paper we present details about the gyrokinetic system and the numerical algorithms used in GX to solve the system. We then present several numerical benchmarks against established gyrokinetic codes in both tokamak and stellarator magnetic geometries to verify that GX correctly simulates gyrokinetic turbulence in the small $\rho _*$. Moreover, we show that the convergence properties of the Laguerre–Hermite spectral velocity formulation are quite favourable for nonlinear problems of interest. Coupled with GPU acceleration, which we also investigate with scaling studies, this enables GX to be able to produce useful turbulence simulations in minutes on one (or a few) GPUs and higher fidelity results in a few hours using several GPUs. GX is open-source software that is ready for fusion reactor design studies.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

CHARMM-GUI Bicelle Builder : An Extension of Membrane Builder for Modeling and Simulation of Bicelle Systems

Membrane mimetics, such as detergent micelles, nanodiscs, and amphipol complexes, which can provide membrane-like environments while retaining small and soluble features, have been utilized to study membrane proteins. A bicelle, composed of varying lipids and detergents, is a useful membrane mimetic because the lipid-to-detergent ratio, the q-value, can be adjusted to alter the properties of the aggregate, including the thickness and size of the bicelle. However, building a bicelle model for modeling and simulation studies requires nontrivial efforts, even for experts. We introduce CHARMM-GUI Bicelle Builder, a web-based platform that can generate various all-atom bicelle systems via a graphical user interface with all available lipids and detergents in Membrane Builder. To illustrate and validate Bicelle Builder with practical systems, we have modeled and simulated pure bicelles consisting of 1,2-dimyristoyl-sn-glycero-3-phosphocholine (DMPC) lipids with 1,2-dihexanoyl-sn-glycero-3-phosphocholine (C6DHPC) detergents and protein–bicelle complexes, composed of DMPC with C6DHPC, foscholine-10 (FOS10), and lysophosphatidylcholine-12 (LPC12) detergents. Our simulation results indicate that Bicelle Builder can generate reliable and robust bicelle models with and without proteins that retain DMPC bilayer characteristics. Bicelle Builder is expected to help researchers better understand not only bicelles themselves but also atomistic-level structures of protein–bicelle complexes that are often difficult to access through experimental approaches.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nanopolysaccharide Builder: A User-Friendly Tool for Atomistic Models of Polysaccharide-Based Nanostructures

Here, we introduce Nanopolysaccharide Builder (NPB), a user-friendly software tool designed to construct polysaccharide nanostructures─mainly those based on cellulose, chitin, and chitosan─using experimental data or user-defined parameters. NPB enables the generation of cellulose and chitin allomorphs with customizable biochemical topologies and also facilitates the construction of large bundles that replicate nanostructures found in biological support systems, including plant cell walls and arthropod cuticles. The software outputs atomic Cartesian coordinates in Protein Data Bank (PDB) format and also provides atom connectivity files in PSF and PARM formats, ensuring seamless integration with major molecular dynamics (MD) engines such as NAMD, CHARMM, GROMACS, AMBER, OpenMM, and LAMMPS. Built on an interactive visualization framework, NPB features a graphical user interface (GUI) and supports both macOS and Linux operating systems. By enabling detailed atomic-scale studies of polysaccharide evolution in extracellular matrices and cell walls of algae, bacteria, fungi, and plants, NPB is poised to advance AI-guided research in sustainable chemical development and biomass utilization.

Wan, Zhangmin [Univ. of British Columbia, Vancouve↗

Speeding Up Hartree–Fock in JuliaChem with Density Fitting

In this work, the density fitting (DF) approximation is added to the restricted Hartree–Fock (RHF) implementation in the JuliaChem computational chemistry code. Utilizing a DF algorithm that uses symmetry and integral screening, a significant reduction in time to compute the Fock matrix is achieved. The symmetry and screening DF-RHF techniques were adapted to be performed on graphics processing units (GPUs), which are well suited to perform the matrix multiplications that comprise the bulk of the Fock build time in DF-RHF. The JuliaChem DF-RHF GPU algorithm employs a novel approach that automatically switches between two DF-RHF algorithms depending on the number of basis functions in the calculation. The JuliaChem GPU DF-RHF implementation demonstrates up to 2× speedup for Fock build times compared to the existing best-in-class GPU DF-RHF implementation by operating directly on screened intermediate matrices. Due to the high portability of the Julia language code, the JuliaChem CPU and GPU DF-RHF implementations could be benchmarked on a variety of CPU and GPU architectures from multiple hardware vendors.

Hayes, John J. [Ames Laboratory, and Iowa State Un↗

COLUMBUS─An Efficient and General Program Package for Ground and Excited State Computations Including Spin–Orbit Couplings and Dynamics

The COLUMBUS program system provides the tools for performing high-level multireference (MR) computations, including the multireference configuration interaction (MRCI) method and its multireference averaged quadratic coupled cluster (MR-AQCC) extension, allowing computations on a wide range of fascinating atomic and molecular systems, including the treatment of open-shells and complicated excited state phenomena. The inclusion of spin−orbit coupling (SOC) directly within the MRCI step enables the description of systems containing heavy elements, such as lanthanides and actinides, whose properties are strongly influenced by SOC. Analytic energy gradients and nonadiabatic couplings at the correlated MRCI level provide the foundation for a variety of dynamics studies, giving insight into ultrafast photochemistry. New and ongoing method developments in COLUMBUS include the computation of spin densities, improved descriptions of ionic states, enhancements to the AQCC method, and the porting of COLUMBUS to graphical processing units (GPUs). New external interfaces enable an enhanced description of electronic resonances and molecules in strong laser fields. This work highlights these new developments while providing a detailed account of the diverse applications of COLUMBUS in recent years.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Implementation of McMurchie–Davidson Algorithm for Gaussian AO Integrals Suited for SIMD Processors

We report an implementation of the McMurchie− Davidson evaluation scheme for 1- and 2-particle Gaussian AO integrals designed for processors with Single Instruction Multiple Data (SIMD) instruction sets. Like in our recent MD implementation for graphical processing units (GPUs) [Asadchev, A.; Valeev, E. F.. J. Chem. Phys. 2024, 160, 244109.], variable-sized batches of shellsets of integrals are evaluated at a time. By optimizing for the floating point instruction throughput rather than minimizing the number of operations, this approach achieves up to 50% of the theoretical hardware peak FP64 performance for many common SIMD-equipped platforms (AVX2, AVX512, NEON), which translates to speedups of up to 30 over the state-of-the-art one-shellset-at-a-time implementation of Obara−Saika-type schemes in Libint for a variety of primitive and contracted integrals. As with our previous work, we rely on the standard C++ programming language such as the std::simd standard library feature to be included in the 2026 ISO C++ standard without any explicit code generation to keep the code base small and portable. The implementation is part of the open source LibintX library freely available at https://github.com/ValeevGroup/libintx.

Basis sets↗

Understanding Solvent-Induced Glass Transition in Polymer Thin Films Using Absorption–Desorption Isotherms

The fundamental thermodynamic and mechanical underpinnings of polymer thin films exposed to solvent vapor are critical for the development of advanced nanolithography and high-performance coatings. This work investigates the solvent− polymer interactions of glassy thin films by using the solvent absorption−desorption isotherms. An analogous relationship to the Flory−Fox equation was observed between solvent−induced glass transition, swelling, Flory−Huggins interaction parameter, and molecular weight. Isothermal swelling measurements revealed that the glass transition trends are more robust in the absorption curve compared to desorption, contrary to previous reports. Excess osmotic pressure analysis of the isotherm provides a measure of the degree of physical aging in thin films annealed below the glass transition. This is further validated in the ordering of block copolymer (BCP) films annealed at low solvent activity. In agreement with the thermal analysis, free-surface plasticization effects become the most prominent below 100 nm. However, solvent annealing is largely dependent on solvent mass transport, as made evident by the strong dependence on solvent viscosity. From these observations, four general types of isotherms are identified that graphically capture distinct solvent−polymer interaction regimes. More broadly, these results inform solvent vapor annealing-induced self-assembly, sequential infiltration synthesis, membrane-based separations, adsorptive processes, and swelling-based responsive materials design.

Hendeniya, Nayanathara [Iowa State Univ., Ames, IA↗

Bridge Connectivity Dictates Spin Interactions and Triplet Pair Dynamics in Intramolecular Singlet Fission

Electron spin plays a critical role in determining the structure, dynamics, and reactivities of molecular excited states, including multiexciton processes such as singlet fission. These systems exhibit triplet pair states whose excited state dynamics can be widely tuned through molecular engineering. For example, the electronic coupling between covalently linked chromophores can be readily modulated using chemical bridges to control proximity, quantum interference, or resonance effects. However, less is known about how spin coupling interactions are impacted by chromophore architecture, and how this influences triplet pair recombination dynamics. Here, in this study, we investigate the role of bridge connectivity and chromophore identity in modulating interchromophore exchange and dipolar coupling interactions for a series of pentacene and tetracene dimers bridged by alternant hydrocarbons (phenylene, naphthalene, anthracene). Using both time-resolved electron paramagnetic resonance and transient absorption spectroscopy, we find that the boundedness and recombination pathways of the triplet pair spins are highly sensitive to molecular architecture and chromophore-specific magnetic dipolar interactions. Notably, nominally ferromagnetic and antiferromagnetic eigenstates result in distinct spin state orderings, consistent with predictions from quantum interference-based graphical models. These findings establish new design principles for tuning spin dynamics in iSF materials, with implications for photonic and quantum information applications.

He, Guiying [City Univ. of New York (CUNY), NY (Un↗

fluxfinder: An R Package for Reproducible Calculation and Initial Processing of Greenhouse Gas Fluxes From Static Chamber Measurements

Fluxes of greenhouse gases are a critical component of the earth's natural climate, but anthropogenic emissions have created an imbalance and resulted in global climate change. Quantifying the emission of these gases is vital to our understanding of their sources and sinks, both natural and anthropogenic. The static chamber method, in which a system of interest is enclosed, and gas concentrations are measured over time, is widely used to estimate fluxes of greenhouse gases. With the development of instruments such as infrared gas analyzers (IRGAs) supporting high-frequency concentration data, there is a growing need for open-source workflows to calculate fluxes. Here we present fluxfinder, an R package designed to support reproducible calculations and processing of greenhouse gas fluxes measured with the static chamber method. The package includes raw data file parsing from widely used IRGAs, metadata matching, unit conversion, flux estimations, and initial quality assurance/quality control (QA/QC). Diagnostic graphical plots provide a transparent way to differentiate between measurement issues and nonlinear behavior. The package is also designed to be easily integrated with the gasfluxes package for further fitting of nonlinear concentration-time models, allowing alternative or additional flux QA/QC. The fluxfinder package offers a flexible workflow that is easily adaptable to promote open and reproducible greenhouse gas flux estimations.

Wilson, Stephanie J.↗

BioRT‐HBV 1.0: A Biogeochemical Reactive Transport Model at the Watershed Scale

Abstract Reactive Transport Models (RTMs) are essential tools for understanding and predicting intertwined ecohydrological and biogeochemical processes on land and in rivers. While traditional RTMs have focused primarily on subsurface processes, recent watershed‐scale RTMs have integrated ecohydrological and biogeochemical interactions between surface and subsurface. These emergent, watershed‐scale RTMs are often spatially explicit and require extensive data, computational power, and computational expertise. There is however a pressing need to create parsimonious models that require minimal data and are accessible to scientists with limited computational background. To that end, we have developed BioRT‐HBV 1.0, a watershed‐scale, hydro‐biogeochemical RTM that builds upon the widely used, bucket‐type HBV model known for its simplicity and minimal data requirements. BioRT‐HBV uses the conceptual structure and hydrology output of HBV to simulate processes including advective solute transport and biogeochemical reactions that depend on reaction thermodynamics and kinetics. These reactions include, for example, chemical weathering, soil respiration, and nutrient transformation. The model uses time series of weather (air temperature, precipitation, and potential evapotranspiration) and initial biogeochemical conditions of subsurface water, soils, and rocks as input, and output times series of reaction rates and solute concentrations in subsurface waters and rivers. This paper presents the model structure and governing equations and demonstrates its utility with examples simulating carbon and nitrogen processes in a headwater catchment. As shown in the examples, BioRT‐HBV can be used to illuminate the dynamics of biogeochemical reactions in the invisible, arduous‐to‐measure subsurface, and their influence on the observed stream or river chemistry and solute export. With its parsimonious structure and easy‐to‐use graphical user interface, BioRT‐HBV can be a useful research tool for users without in‐depth computational training. It can additionally serve as an educational tool that promotes pollination of ideas across disciplines and foster a diverse, equal, and inclusive user community.

Sadayappan, Kayalvizhi↗