Search NASA⌕ Search

SEARCH · Search NASA

Results for “Massive parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

Simulation of 24,000 Electron Dynamics: Real-Time Time-Dependent Density Functional Theory (TDDFT) with the Real-Space Multigrids (RMG)

Here, we present the theory, implementation, and benchmarking of a real-time time-dependent density functional theory (RT-TDDFT) module within the RMG code, designed to simulate the electronic response of molecular systems to external perturbations. Our method offers insights into nonequilibrium dynamics and excited states across a diverse range of systems, from small organic molecules to large metallic nanoparticles. Benchmarking results demonstrate excellent agreement with established TDDFT implementations and showcase the superior stability of our time integration algorithm, enabling long-term simulations with minimal energy drift. The scalability and efficiency of RMG on massively parallel architectures allow for simulations of complex systems, such as plasmonic nanoparticles with thousands of atoms. Future extensions, including nuclear and spin dynamics, will broaden the applicability of this RT-TDDFT implementation, providing a powerful toolset for studies of photoactive materials, nanoscale devices, and other systems where real-time electronic dynamics is essential.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GPU-Accelerated Solution of the Bethe–Salpeter Equation for Large and Heterogeneous Systems

We present a massively parallel GPU-accelerated implementation of the Bethe–Salpeter equation (BSE) for the calculation of the vertical excitation energies (VEEs) and optical absorption spectra of condensed and molecular systems, starting from single-particle eigenvalues and eigenvectors obtained with density functional theory. The algorithms adopted here circumvent the slowly converging sums over empty and occupied states and the inversion of large dielectric matrices through a density matrix perturbation theory approach and a low-rank decomposition of the screened Coulomb interaction, respectively. Further computational savings are achieved by exploiting the nearsightedness of the density matrix of semiconductors and insulators to reduce the number of screened Coulomb integrals. We scale our calculations to thousands of GPUs with a hierarchical loop and data distribution strategy. The efficacy of our method is demonstrated by computing the VEEs of several spin defects in wide-band-gap materials, showing that supercells with up to 1000 atoms are necessary to obtain converged results. We discuss the validity of the common approximation that solves the BSE with truncated sums over empty and occupied states. In conclusion, we then apply our GW-BSE implementation to a diamond lattice with 1727 atoms to study the symmetry breaking of triplet states caused by the interaction of a point defect with an extended line defect.

Absorption spectra↗

Stochastic GW -GPU: Rapid Quasi-Particle Energies for Molecules beyond 10,000 Atoms

StochasticGW is a code for computing accurate quasi-particle (QP) energies of molecules and material systems in the GW approximation. StochasticGW utilizes the stochastic Resolution of the Identity (sROI) technique to enable a massively parallel implementation with computational costs that scale semilinearly with system size, allowing the method to access systems with tens of thousands of electrons. Here, we introduce a new implementation, StochasticGW-GPU, for which the main bottleneck steps have been ported to GPUs and give substantial performance improvements over previous versions of the code. We showcase the new code by computing band gaps of hydrogenated silicon clusters (Si x H y ) containing up to 10,001 atoms and 35,144 electrons, and we obtain individual QP energies with a statistical precision of better than ±0.03 eV with times-to-solution of less than 1 h.

Thomas, Phillip S. [Lawrence Berkeley National Lab↗

Resistive Switching of Spinel Li 4 Ti 5 O 12 Lithium-Ion Battery Material for Neuromorphic Computing

The rapid rise of AI has exposed significant limitations in conventional Von Neumann computing architecture, particularly in regard to speed and energy efficiency. To address these challenges, researchers are exploring a brain-inspired neuromorphic architecture that mimics biological neural networks, enabling massive parallel processing with reduced power consumption for complex AI computational demands. Recent interest has focused on utilizing battery electrodes and solid electrolyte materials for their resistive switching properties in developing a neuromorphic architecture. These properties are precisely tuned through local- and bulk-level chemical composition modifications via voltage bias stimuli. In this study, we demonstrate fabricating a three-terminal lithium-ion electrochemical transistor based on lithium titanium oxide (Li 4 Ti 5 O 12 ), a popular lithium-ion battery anode material. We deposited and characterized LTO thin films using RF sputtering, demonstrating a 6 orders of magnitude increase in electronic conductivity upon lithiation, with conductivity plateauing after 20% lithiation. Density functional theory calculations revealed transformation from the insulating to conducting state, supported by experimental characterization through X-Ray Photoelectron Spectroscopy (XPS) and Direct Current (DC) polarization analyses. The fabricated transistor consisted of LTO as the channel layer, gold as source/drain terminals, lithium phosphorus oxynitride (LiPON) as the lithium-ion conductor, and copper as the gate terminal. The device exhibited clear hysteresis in transfer characteristics due to lithium insertion/extraction processes. Long-term potentiation (LTP) and long-term depression (LTD) measurements showed an asymmetric ratio of 1.425 and maximum/minimum conductance ratio of 7.83. When implemented in a deep neural network (DNN) for MNIST handwritten digit recognition, the device achieved 92.03% accuracy over 20 training epochs. Detailed transport mechanism analysis revealed the crucial role of oxygen vacancies and interface effects in device operation. Our preliminary findings establish LTO-based lithium-ion electrochemical transistors as promising candidates for energy-efficient neuromorphic computing applications, offering potential solutions to traditional Von Neumann architecture limitations.

25 ENERGY STORAGE↗

Identification of functional non-coding variants associated with orofacial cleft

Oral facial cleft (OFC) comprises cleft lip with or without cleft palate (CL/P) or cleft palate only. Genome wide association studies (GWAS) of isolated OFC have identified common single nucleotide polymorphisms (SNPs) in many genomic loci where the presumed effector gene (for example, IRF6 in the 1q32 locus) is expressed in embryonic oral epithelium. To identify candidates for functional SNPs at eight such loci we conduct a massively parallel reporter assay in a fetal oral epithelial cell line, revealing SNPs with allele-specific effects on enhancer activity. We filter these SNPs against chromatin-mark evidence of enhancers and test a subset in traditional reporter assays, which support the candidacy of SNPs at loci containing FOXE1, IRF6, MAFB, TFAP2A, and TP63. For two SNPs near IRF6 and one near FOXE1, we engineer the genome of induced pluripotent stem cells, differentiate the cells into embryonic oral epithelium, and discover allele-specific effects on the levels of effector gene expression, and, in two cases, the binding affinity of transcription factors FOXE1 or ETS2. Conditional analyses of GWAS data suggest the two functional SNPs near IRF6 account for the majority of risk for CL/P at this locus. This study connects genetic variation associated with OFC to mechanisms of pathogenesis.

Kumari, Priyanka↗

A map of the rubisco biochemical landscape

Rubisco is the primary CO 2 -fixing enzyme of the biosphere, yet it has slow kinetics. The roles of evolution and chemical mechanism in constraining its biochemical function remain debated. Engineering efforts aimed at adjusting the biochemical parameters of rubisco have largely failed, although recent results indicate that the functional potential of rubisco has a wider scope than previously known. Here we developed a massively parallel assay, using an engineered Escherichia coli in which enzyme activity is coupled to growth, to systematically map the sequence–function landscape of rubisco. Composite assay of more than 99% of single-amino acid mutants versus CO 2 concentration enabled inference of enzyme velocity and apparent CO 2 affinity parameters for thousands of substitutions. This approach identified many highly conserved positions that tolerate mutation and rare mutations that improve CO 2 affinity. These data indicate that non-trivial biochemical changes are readily accessible and that the functional distance between rubiscos from diverse organisms can be traversed, laying the groundwork for further enzyme engineering efforts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An expanded registry of candidate cis -regulatory elements

Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq, massively parallel reporter assay, CRISPR perturbation and transgenic mouse assays have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.

Moore, Jill E. [Univ. of Massachusetts, Worchester↗

Improving the precision of forces in real-space pseudopotential density functional theory

The high-order finite difference real-space pseudopotential density functional theory (DFT) approach is a valuable method for large-scale, massively parallel DFT calculations. A significant challenge in the approach is the oscillating “egg-box” error introduced by aliasing associated with a coarse grid spacing. To address this issue while minimizing computational cost, we developed a finite difference interpolation (FDI) scheme [Roller et al., J. Chem. Theory Comput. 19, 3889 (2023)] as a means of exploiting the high resolution of the pseudopotential to reduce egg-box effects systematically. Here, we show an implementation of this method in the PARSEC code and examine the practical utility of the combination of FDI with additional methods for improving force precision and/or reducing its computational cost, including orbital-based forces, compensating charges (namely, adding and subtracting a judiciously chosen charge density such that the total density is unaltered), and a modified spatial domain in which the real-space grid is defined. Using selected small molecules, as well as metallic Li, as test cases, we show that a combination of all four aspects leads to a significant reduction in computational cost while retaining a high level of precision that supports accurate structures and vibrational spectra, as well as stable and accurate molecular dynamics runs.

Chemistry↗

Developing reliable machine learning interatomic potential for Fe–Cr–Ni austenitic alloys

Gaining atomistic understanding of mechanical behavior of heat-resistant structural materials such as Fe–Cr–Ni-based alloys requires an approach with an accuracy close to density functional theory (DFT) that considers the intrinsic properties of the bulk lattice and important defects such as stacking faults, grain boundaries, and surfaces. This work aims to develop reliable machine learning interatomic potential (MLIAP) at cross-scale for Fe–Cr–Ni ternary alloys with a focus on the face-centered-cubic (fcc) solid solution structure. Leveraging the advantages of moment tensor potentials, which typically necessitate a relatively small training dataset and enable rapid calculations using the large-scale atomic/molecular massively parallel simulator package, we ensure the stability and accuracy of the trained potentials. Important defects such as stacking faults, grain boundaries, and surfaces for wide-range compositions are investigated. Structural, thermal, elastic, and defect properties are determined from molecular dynamics simulations comprising several thousand atoms, generated via canonical Monte Carlo simulations guided by the trained potential. The trained potential allows efficient atomic simulations of structural, thermal, and mechanical properties of fcc Fe–Cr–Ni solid solution alloys as a function of composition and temperature. Therefore, the MLIAP approach represents a major advancement from DFT calculations that are limited to small simulation sizes and traditional molecular dynamics simulations using relatively low accuracy potentials. Furthermore, this work outlines a practical foundation for further investigating the structural evolution and mechanical behavior of austenitic stainless steel and nickel-based alloys in a wide array of applications in extreme environments.

Crystal structure↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Positive Neutrino Masses with DESI DR2 via Matter Conversion to Dark Energy

The Dark Energy Spectroscopic Instrument (DESI) is a massively parallel spectroscopic survey on the Mayall telescope at Kitt Peak, which has released measurements of baryon acoustic oscillations determined from over 14 million extragalactic targets. We combine DESI Data Release 2 with CMB datasets to search for evidence of matter conversion to dark energy (DE), focusing on a scenario mediated by stellar collapse to cosmologically coupled black holes (CCBHs). In this physical model, which has the same number of free parameters as Λ⁢CDM, DE production is determined by the cosmic star formation rate density (SFRD), allowing for distinct early- and late-time cosmologies. Using two SFRDs to bracket current observations, we find that the CCBH model: accurately recovers the cosmological expansion history, agrees with early-time baryon abundance measured by BBN, reduces tension with the local distance ladder, and relaxes constraints on the summed neutrino mass ∑𝑚 𝜈 . For these SFRDs, we find a peaked positive ∑𝑚 𝜈 < 0.149 eV (95% confidence) and ∑𝑚 𝜈 = 0.106$^{+0.050}_{−0.069}$ eV, respectively, in good agreement with lower limits from neutrino oscillation experiments. A peak in ∑𝑚 𝜈 > 0 results from late-time baryon consumption in the CCBH scenario and is expected to be a general feature of any model that converts sufficient matter to dark energy during and after reionization.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Integrating machine learning interatomic potentials with hybrid reverse Monte Carlo structure refinements in RMCProfile

Structure refinement with reverse Monte Carlo (RMC) is a powerful tool for interpreting experimental diffraction data. To ensure that the under-constrained RMC algorithm yields reasonable results, the hybrid RMC approach applies interatomic potentials to obtain solutions that are both physically sensible and in agreement with experiment. To expand the range of materials that can be studied with hybrid RMC, we have implemented a new interatomic potential constraint in RMCProfile that grants flexibility to apply potentials supported by the Large-scale Atomic/Molecular Massively Parallel Simulator ( LAMMPS ) molecular dynamics code. This includes machine learning interatomic potentials, which provide a pathway to applying hybrid RMC to materials without currently available interatomic potentials. To this end, we present a methodology to use RMC to train machine learning interatomic potentials for hybrid RMC applications.

Cuillier, Paul↗

Femtojoule optical nonlinearity for deep learning with incoherent illumination

Optical neural networks (ONNs) are a promising computational alternative for deep learning due to their inherent massive parallelism for linear operations. However, the development of energy-efficient and highly parallel optical nonlinearities, a critical component in ONNs, remains an outstanding challenge. Here, we introduce a nonlinear optical microdevice array (NOMA) compatible with incoherent illumination by integrating the liquid crystal cell with silicon photodiodes at the single-pixel level. We fabricate NOMA with more than half a million pixels, each functioning as an optical analog of the rectified linear unit at ultralow switching energy down to 100 femtojoules per pixel. With NOMA, we demonstrate an optical multilayer neural network. Our work holds promise for large-scale and low-power deep ONNs, computer vision, and real-time optical image processing.

36 MATERIALS SCIENCE↗

High-entropy 1D halide perovskite piezoelectrics found by megalibrary synthesis and rapid nonlinear optical screening

Piezoelectric molecular crystals offer excellent compositional and structural tunability and sustainable processability. However, their discovery is slow, primarily due to the serial synthesis and screening processes used. Here, we report an approach that combines massively parallel megalibrary synthesis with scanning second harmonic generation (SHG) microscopy for rapid screening of piezoelectric molecular crystals. Megalibraries consisting of more than 1,000,000 compositionally distinct but positionally encoded TMCM x TMA (1–x) Cd y Pb (1–y) ClzBr (3–z) (TMCM: trimethylchloromethylammonium, TMA: tetramethylammonium; 0 ≤ x ≤ 1, 0 ≤ y ≤ 1, 0 ≤ z ≤ 3) nanocrystals were synthesized. The megalibraries were rapidly screened by SHG microscopy to identify notable noncentrosymmetric structures, which were then tested for piezoelectricity, facilitating discovery of a high-entropy noncentrosymmetric material with a large d 33 (TMCM 0.75 TMA 0.25 Cd 0.75 Pb 0.25 Cl 1.5 Br 1.5 , 42.8 picocoulombs per newton). Furthermore, this approach enabled systematic investigation of the Curie temperature (T C )–composition relationship in the TMCMCdCl z Br (3–z) system, facilitating reverse design of materials with targeted T C . Our work establishes a powerful approach to accelerate the discovery and design of unusual piezoelectrics for next-generation electronics and optics.

Li, Jun [Northwestern University, Evanston, IL (Un↗

TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU Systems

Deep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation. The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability. With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy. Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference. However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies. While an alternative solution is to test on GPU simulators, they are often too slow for these l

Li, Ying [William & Mary, Williamsburg, VA, USA] (↗

ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling

Sparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data.

Wang, Xiao [ORNL] (ORCID:0000000165451943)↗

Vidyut3d: A Non-Equilibrium Plasma Modeling Tool [SWR-24-101]

Vidyut3d is a massively-parallel plasma-fluid solver for low-temperature plasmas (LTPs) that supports both local field (LFA) and local mean energy (LMEA) approximations, as well as complex gas and surface-phase chemistry. The solver supports 2D and 3D domains, and uses AMReX's adaptive mesh refinement capabilities to increase the grid resolution around complex structures (e.g. streamer heads and sheaths) while maintaining a tractable problem size. Vidyut specializes in simulating various types of gas-phase discharges, as well as plasma/surface interactions and surface chemistry (e.g. for plasma-mediated catalysis applications). The solver also supports hybrid CPU/GPU parallelization strategies, and has demonstrated excellent scaling on various HPC architectures for problem sizes consisting of O(100 M) control volumes.

Sitaraman, Hariswaran↗