Search NASA⌕ Search

SEARCH · Search NASA

Results for “RAJA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

EQSIM and RAJA: Enabling Exascale Predictions of Earthquake Effects on Critical Infrastructure

Nearly 120 years ago, the great “San Francisco” earthquake of 1906 provided a stark and sobering view of the havoc that can be caused by the sudden and violent movement of Earth’s tectonic plates. According to USGS, the rupture along the San Andreas fault extended 296 miles (447 kilometers) and shook so violently that the motion could be felt as far north as Oregon and east into Nevada. The estimated 7.9-magnitude quake and subsequent fires decimated the major metropolis and surrounding areas: buildings turned to ruins, hundreds of thousands of people left homeless, and a death toll exceeding 3,000. Today, as evidenced by the catastrophic 7.8-magnitude earthquake that struck Turkey in February of 2023, large earthquakes still present a significant potential danger to life and economic security as researchers work to develop ways to better understand earthquake phenomena and quantify associated risks.

42 ENGINEERING↗

Toward performance-portable PETSc for GPU-based exascale systems

The Portable Extensible Toolkit for Scientific computation (PETSc) library delivers scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization. The PETSc design for performance portability addresses fundamental GPU accelerator challenges and stresses flexibility and extensibility by separating the programming model used by the application from that used by the library, and it enables application developers to use their preferred programming model, such as Kokkos, RAJA, SYCL, HIP, CUDA, or OpenCL, on upcoming exascale systems. Furthermore, a blueprint for using GPUs from PETSc-based codes is provided, and case studies emphasize the flexibility and high performance achieved on current GPU-based systems.

97 MATHEMATICS AND COMPUTING↗

Performance Portable Graphics Processing Unit Acceleration of a High-Order Finite Element Multiphysics Application

The Lawrence Livermore National Laboratory (LLNL) will soon have in place the El Capitan exascale supercomputer, based on advanced micro devices (AMD) graphics processing units (GPUs). As part of a multiyear effort under the National Nuclear Security Administration (NNSA) Advanced Simulation and Computing (ASC) program, we have been developing marbl, a next generation, performance portable multiphysics application based on high-order finite elements. In previous years, we successfully ported the Arbitrary Lagrangian–Eulerian (ALE), multimaterial, compressible flow capabilities of marbl to nvidia GPUs as described in Vargas et al. Here, in this paper, we describe our ongoing effort in extending marbl's GPU capabilities with additional physics, including multigroup radiation diffusion and thermonuclear burn for high energy density physics (HEDP) and fusion modeling. We also describe how our portability abstraction approach based on the raja Portability Suite and the mfem finite element discretization library has enabled us to achieve high performance on AMD based GPUs with minimal effort in hardware-specific porting. Throughout this work, we highlight numerical and algorithmic developments that were required to achieve GPU performance.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

SW4 Curvilinear Kernels

Five computationally expensive stencil evaluation routines from SW4(https://github.com/geodynamics/sw4 GPL license) have been extracted and packaged with a driver to create a mini-app for evaluating compiler and GPU performance. The kernels can executed on AMD and Nvidia GPUs with and without RAJA. The kernel driver generates synthetic inputs, checks for correctness and measures kernel run times.

Pankajakshan, Ramesh↗

Porting the Nonlinear Optimization Library HiOp to Accelerator-Based Hardware Architectures

While interior point method has been the centerpiece of nonlinear programming tools used in science and engineering, its reliance on linear solvers that can tackle sparse symmetric indefinite and highly ill-conditioned problems made it difficult to implement it effectively on hardware accelerators. HiOp optimization package attempts to provide an implementation of the interior point method suitable for hardware accelerators by compressing the original sparse problem to produce an underlying linear problem that is dense and of manageable size. Implementations of dense linear solvers are more mature and utilize hardware accelerators better than their sparse counterparts. There is a number of important domain problems, such as optimal power flow analysis for power grids, where the sparse problem can be effectively compressed and deploying dense linear solver within the interior point method can improve performance. Here we describe a portable implementation of HiOp optimization engine, which uses a linear solver from Magma library and runs entirely on hardware accelerators. To compress the problem, HiOp uses customized mixed dense-sparse linear algebra. All HiOp kernels are implemented using Umpire and RAJA portability libraries. We describe details of the implementation and discuss trade-offs between performance, portability and development cost.

97 MATHEMATICS AND COMPUTING↗

Giant low-field magnetocaloric effect and refrigerant capacity in reduced dimensionality EuTiO 3 multiferroics

Engineering magnetic materials into a thin film form while preserving its excellent magnetocaloric response is essential in the development of miniature magnetic coolers. We demonstrate how this can be achieved in the case of EuTiO 3 - an emerging multiferroic material. Unlike conventional cases where reduced dimensionality considerably decreased the magnetic entropy change (ΔS M ) and hence the refrigerant capacity (RC), we show the large low-field enhancements of ΔS M and RC in a ~100 nm thick nanocrystalline EuTiO 3 film (ΔS M ~ 24 J kg -1 K -1 and RC = 152 J kg -1 for μ 0 ΔH = 2T) relative to its single crystal counterpart (ΔS M ~ 17 J kg -1 K -1 and RC ~ 107 J kg -1 for μ 0 ΔH = 2T). The nanocrystalline EuTiO 3 film is an excellent candidate for cryogenic magnetic refrigeration. From our study, a new approach for improving both MCE and RC in magnetic nanomaterials is proposed, which will stimulate further research on magnetocaloric thin films and related cooling devices.

36 MATERIALS SCIENCE↗

Evaluation of residual stresses in isothermal friction stir welded 304L stainless steel plates

Friction stir welding was performed on 304L SS plates in order to heal simulated cracks created by electrical discharge machining. Two different tool temperatures (825 and 725 °C) were chosen for this study. Both neutron diffraction and X-ray diffraction techniques were employed to evaluate the residual stresses along two orthogonal reference directions, longitudinal (syy) and transverse (sxx). The former technique was also used to measure residual stresses at various depths. It was found that, at 1 mm depth from the top surface inside the stir zone (SZ), the longitudinal component was tensile in nature while the transverse component was compressive. The nature and magnitude of the residual stress fields, and the position of the peak residual stresses were found to vary with the weld depth. The SZ of the 725 °C weld exhibited higher peak stress than 825 °C weld mainly due to a lack of stress relief at the lower temperature.

Friction stir welding, Steel↗

Particle-in-cell Monte Carlo-collision modeling of non-ideal effects in wave-heated dense microplasmas

A computational model for non-ideal plasma effects during the time evolution of a second-stage laser-heated discharge at high pressures is presented. The model extends a classical one-dimensional particle-in-cell Monte Carlo-collision (PIC-MCC) approach coupled with Maxwell's equations for the laser-heating process of a xenon plasma at 300 K temperature and 10–100 bar pressure. Plasma non-ideality resulting from Coulomb coupling at high plasma densities is manifested as a depression in the effective ionization potential of atoms and enhanced collision cross sections. These non-ideal effects are represented using the Ecker–Kröll model in the context of the PIC-MCC approach. We find that full ionization of the plasma is obtained on the picosecond timescale, starting from the skin layer and quickly expanding throughout the domain through an anomalous extension of the skin depth. More critically, we show that the inclusion of the non-ideal plasma effects results in more rapid ionization compared to an ideal plasma, especially at higher pressures. The ionization delay reduction is of the order of a fraction of a picosecond, corresponding to a 16% decrease at 100 bar. As the amplitude of the wave field is lowered, the ionization rate is lowered, making the plasma non-ideality effects more prominent.

Solmaz, Evrim (ORCID:0000000211830881)↗

Kramers nodal line in the charge density wave state of YTe 3 and the influence of twin domains

Recent studies have focused on the relationship between charge density wave (CDW) collective electronic ground states and nontrivial topological states. YTe 3 , a nonmagnetic quasi-two-dimensional chalcogenide, has been reported to exhibit a CDW state below 334 K. Using angle-resolved photoemission spectroscopy (ARPES) and density functional theory (DFT), we establish that YTe 3 is a CDW-induced Kramers nodal line (KNL) metal, a recently proposed topological state of matter. Scanning tunneling microscopy and low energy electron diffraction reveal two orthogonal domains, each with a unidirectional CDW and a similar wave vector (𝐪 CDW ). When the influence of twin domains is considered, the effective band structure (EBS) computations that utilize DFT-calculated bands using a noncentrosymmetric structure determined by x-ray crystallography, show excellent agreement with ARPES. The noncentrosymmetry of YTe 3 is established by Raman spectroscopy. The Fermi surface and ARPES intensity plots show weak shadow bands displaced by 𝐪 CDW from the main bands. Furthermore, these are linked to CDW modulation, as the EBS calculation confirms. Bilayer split main and shadow bands suggest the existence of crossings, according to theory and experiment. DFT bands, including spin-orbit coupling, indicate existence of a KNL along the Σ direction from multiple crossings of bands dispersing perpendicular to it. Additionally, doubly degenerate bands are only found along the KNL at all energies, with some bands dispersing through the Fermi level.

Angle-resolved photoemission spectroscopy↗

Physics-informed Deep Reinforcement Learning-based Control in Power systems

Incorporating physics information into the deep reinforcement learning (DRL) process is a promising approach for addressing the challenges faced in learning-based control design problems for physical systems. Power grid dynamics, being a physical system, adheres to specific physical laws, constraints, as well as operational and control rules. Therefore, consideration of such physics-based law improves the learning process drastically. In general, traditional grid control schemes rely on rule-based mechanisms that cannot adapt to changing operating conditions. To improve the adaptability and computation time, recent research has seen a surge of DRL-based applications in power grid control. A generic DRL-based control design imposes the system performance requirements through the design of reward functions. In some cases, some of the important physics information is injected through this reward function. However, due to the complex dynamics and large state-action space, learning an optimal DRL policy often becomes challenging. Inspired by the latest developments in general machine learning (ML) research, power system researchers have been investigating more direct ways of incorporating physics knowledge into DRL training. This chapter specifically focuses on these aspects of physics-informed DRL designs in grid control. It discusses the significance, applications, research gaps, and open problems that need to be addressed in future research.

artificial intelligence, machine learning↗

Simultaneously Enhancing Exciton/Charge Transport in Organic Solar Cells by an Organoboron Additive

Efficient exciton diffusion and charge transport play a vital role in advancing the power conversion efficiency (PCE) of organic solar cells (OSCs). Here, a facile strategy is presented to simultaneously enhance exciton/charge transport of the widely studied PM6:Y6-based OSCs by employing highly emissive trans-bis(dimesitylboron)stilbene (BBS) as a solid additive. BBS transforms the emissive sites from a more H-type aggregate into a more J-type aggregate, which benefits the resonance energy transfer for PM6 exciton diffusion and energy transfer from PM6 to Y6. Transient gated photoluminescence spectroscopy measurements indicate that addition of BBS improves the exciton diffusion coefficient of PM6 and the dissociation of PM6 excitons in the PM6:Y6:BBS film. Transient absorption spectroscopy measurements confirm faster charge generation in PM6:Y6:BBS. Moreover, BBS helps improve Y6 crystallization, and current-sensing atomic force microscopy characterization reveals an improved charge-carrier diffusion length in PM6:Y6:BBS. Finally, owing to the enhanced exciton diffusion, exciton dissociation, charge generation, and charge transport, as well as reduced charge recombination and energy loss, a higher PCE of 17.6% with simultaneously improved open-circuit voltage, short-circuit current density, and fill factor is achieved for the PM6:Y6:BBS devices compared to the devices without BBS (16.2%).

14 SOLAR ENERGY↗

Quasi‐Homojunction Organic Nonfullerene Photovoltaics Featuring Fundamentals Distinct from Bulk Heterojunctions

Abstract In contrast to classical bulk heterojunction (BHJ) in organic solar cells (OSCs), the quasi‐homojunction (QHJ) with extremely low donor content (≤10 wt.%) is unusual and generally yields much lower device efficiency. Here, representative polymer donors and nonfullerene acceptors are selected to fabricate QHJ OSCs, and a complete picture for the operation mechanisms of high‐efficiency QHJ devices is illustrated. PTB7‐Th:Y6 QHJ devices at donor:acceptor (D:A) ratios of 1:8 or 1:20 can achieve 95% or 64% of the efficiency obtained from its BHJ counterpart at the optimal D:A ratio of 1:1.2, respectively, whereas QHJ devices with other donors or acceptors suffer from rapid roll‐off of efficiency when the donors are diluted. Through device physics and photophysics analyses, it is observed that a large portion of free charges can be intrinsically generated in the neat Y6 domains rather than at the D/A interface. Y6 also serves as an ambipolar transport channel, so that hole transport as also mainly through Y6 phase. The key role of PTB7‐Th is primarily to reduce charge recombination, likely assisted by enhancing quadrupolar fields within Y6 itself, rather than the previously thought principal roles of light absorption, exciton splitting, and hole transport.

Chemistry↗

Size‐Induced Ferroelectricity in Antiferroelectric Oxide Membranes

Abstract Despite extensive studies on size effects in ferroelectrics, how structures and properties evolve in antiferroelectrics with reduced dimensions still remains elusive. Given the enormous potential of utilizing antiferroelectrics for high‐energy‐density storage applications, understanding their size effects will provide key information for optimizing device performances at small scales. Here, the fundamental intrinsic size dependence of antiferroelectricity in lead‐free NaNbO 3 membranes is investigated. Via a wide range of experimental and theoretical approaches, an intriguing antiferroelectric‐to‐ferroelectric transition upon reducing membrane thickness is probed. This size effect leads to a ferroelectric single‐phase below 40 nm, as well as a mixed‐phase state with ferroelectric and antiferroelectric orders coexisting above this critical thickness. Furthermore, it is shown that the antiferroelectric and ferroelectric orders are electrically switchable. First‐principle calculations further reveal that the observed transition is driven by the structural distortion arising from the membrane surface. This work provides direct experimental evidence for intrinsic size‐driven scaling in antiferroelectrics and demonstrates enormous potential of utilizing size effects to drive emergent properties in environmentally benign lead‐free oxides with the membrane platform.

36 MATERIALS SCIENCE↗

Size-Induced Ferroelectricity in Antiferroelectric Oxide Membranes (Adv. Mater. 17/2023)

Thin Films. In article number 2210562, Ruijuan Xu, Kevin J. Crust, Varun Harbola, and co-workers report intrinsic size-driven scaling in lead-free antiferroelectric thin films. They demonstrate an intriguing antiferroelectric-to-ferroelectric transition upon reducing the thickness of antiferroelectric NaNbO 3 membranes. Here, the image shows the coexistence of ferroelectric and antiferroelectric phases in freestanding NaNbO 3 membranes.

36 MATERIALS SCIENCE↗

Facile Tensile Testing Platform for In Situ Transmission Electron Microscopy of Nanomaterials

In situ tensile testing using transmission electron microscopy (TEM) is a powerful technique to probe structure-property relationships of materials at the atomic scale. In this work, a facile tensile testing platform for in situ characterization of materials inside a transmission electron microscope is demonstrated. The platform consists of: 1) a commercially available, flexible, electron-transparent substrate (e.g., TEM grid) integrated with a conventional tensile testing holder, and 2) a finite element simulation providing quantification of specimen-applied strain. The flexible substrate (carbon support film of the TEM grid) mitigates strain concentrations usually found in free-standing films and enables in situ straining experiments to be performed on materials that cannot undergo localized thinning or focused ion beam lift-out. The finite element simulation enables direct correlation of holder displacement with sample strain, providing upper and lower bounds of expected strain across the substrate. The tensile testing platform is validated for three disparate material systems: sputtered gold-palladium, few-layer transferred tungsten disulfide, and electrodeposited lithium, by measuring lattice strain from experimentally recorded electron diffraction data. The results show good agreement between experiment and simulation, providing confidence in the ability to transfer strain from holder to sample and relate TEM crystal structural observations with material mechanical properties.

2D materials↗