Search NASASearch

SEARCH · Search NASA

Results for “Edge computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Establish the basis for Breadth-First Search on Frontier System: XBFS on AMD GPUs

Graphics Processing Units (GPUs) offer significant potential for accelerating various computational tasks, including Breadth-First Search (BFS). Numerous efforts have been made to deploy BFS on GPUs effectively. To address the dynamic nature of BFS, XBFS, the state-of-the-art work, employs an adaptive strategy that leverages different optimized frontier queue generation designs, accommodating the varying characteristics of levels in BFS. While XBFS demonstrates excellent performance on NVIDIA Quadro P6000 GPUs, it faces challenges when deployed on AMD GPUs. In this work, we present our efforts to implement XBFS’s adaptive approach on Frontier, the most powerful supercomputer system, by porting XBFS to AMD MI250X GPUs. Through targeted optimizations tailored to the unique features of AMD GPUs, our implementation achieves an average performance of 43 Giga-Traversed Edges Per Second (GTEPS) per Graphics Compute Dies (GCD). Based on these results, we observe potential for surpassing the performance of the official Frontier results from the Graph500 benchmark released in June 2024.

Yang, Haoshen

A constrained-transport embedded boundary method for compressible resistive magnetohydrodynamics

Motivated by the increased interest in pulsed-power magneto-inertial fusion devices in recent years, we present a method for implementing an arbitrarily shaped embedded boundary on a Cartesian mesh while solving the equations of compressible resistive magnetohydrodynamics. The method is built around a finite volume formulation of the equations in which a Riemann solver is used to compute fluxes on the faces between grid cells, and a face-centered constrained transport formulation of the induction equation. The small time step problem associated with the cut cells is avoided by always computing fluxes on the faces and edges of the Cartesian mesh. We extend the method to model a moving interface between two materials with different properties using a ghost-fluid approach, and show some preliminary results including shock-wave-driven and magnetically-driven dynamical compressions of magnetohydrostatic equilibria. In conclusion, we present a thorough verification of the method and show that it converges at second order in the absence of discontinuities, and at first order with a discontinuity in material properties.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Dendritic Computing with Multigate Ferroelectric Field-Effect Transistors

Although inspired by neuronal systems in the brain, artificial neural networks generally employ point-neurons, which offer computational complexity far less than that of their biological counterparts. Neurons have dendritic arbors that connect to different sets of synapses and offer local nonlinear accumulation – playing a pivotal role in processing and learning. Inspired by this, we propose a novel neuron design based on a multigate ferroelectric field-effect transistor that mimics dendrites. It leverages ferroelectric nonlinearity for local computations within dendritic branches while utilizing the transistor action to generate the neuronal output. The branched architecture enables smaller crossbar arrays in hardware integration, improving efficiency. Using an experimentally calibrated device-circuit-algorithm cosimulation framework, we demonstrate that networks incorporating our dendritic neurons achieve superior performance compared to much larger networks without dendrites (∼ 17× fewer trainable weight parameters). These findings suggest that dendritic hardware can significantly improve computational efficiency and learning capacity of neuromorphic systems optimized for edge applications.

36 MATERIALS SCIENCE

Towards Cross-Facility Workflows Orchestration through Distributed Automation

Modern science relies on end-to-end workflows that incorporate experimental instruments and utilize edge, cloud, or high-performance computing and storage resources. These components are geographically dispersed across various user facilities and interconnected through high-speed networks. In this paper, we present Zambeze, an automated distributed framework designed to facilitate this new class of cross-facility workflows. Utilizing swarm intelligence principles, Zambeze orchestrates science campaigns by managing distributed autonomous agents. These agents can offer a suite of services, including computing, storage, and data management. We demonstrate the feasibility of Zambeze through a real-world application involving electron microscopy, enhanced with Artificial Intelligence capabilities.

Skluzacek, Tyler

Basic Physical Processes Involving Dust in Fusion Plasmas

This report presents the main outcomes of our research focused on understanding the behavior and effects of dust particles in fusion plasmas, particularly in the edge regions of tokamaks and stellarators. We conducted computer modeling studies of burst injections of carbon and tungsten dust particles in DIII-D like divertor plasmas. The studies investigated effects of transient influx of the low- and high-Z material plasma contaminants in form of dust on the edge plasma dynamics in a modern mid-size tokamak. We also performed theoretical and computational studies of the forces acting on non-spherical dust grains in magnetized and nonmagnetized plasmas. In addition, we cooperated with General Atomics slag management group on development of DIII-D dust injection experiments to measure trajectories of dust grains in the divertor plasma and compare them with DUSTT code predictions for validation of the dust modeling capabilities. We collaborated with JET experimentalists on evaluation of radiative losses induced by injection of mixed neon-deuterium ice pellets for disruption and run-away electron generation mitigation during thermal and current quench phases. We also cooperated with experimentalists at LHD fusion device (Japan) on modeling support of experiments on injection of boron granules in fusion plasma discharges to assess their dynamics and effectiveness for in situ dynamic wall conditioning.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

From Reproducible Edge–Cloud Experimentation to Real-World Practice: The E2Clab Experience

Reproducibility is already difficult in distributed systems; on the computing continuum, it becomes substantially harder. Applications that span sensing devices, edge and fog resources, and cloud platforms must be evaluated across heterogeneous hardware, variable network conditions, cross-layer orchestration decisions, and long-running workflow lifecycles. We use E2Clab as a case study to examine these challenges and their implications for experimental methodology. We explain why reproducible experimentation is harder on the continuum, then revisit E2Clab as an initial response based on explicit modeling of infrastructure, workflow lifecycle, and artifacts. Lastly, we discuss how its evolution toward more realistic application settings can be understood through the lens of Translational Computer Science. We argue that reproducible continuum experimentation requires methods that are rigorous enough for research while remaining adaptable to real-world practice.

42 ENGINEERING

Buried Dirac Points in Quantum Spin Hall Insulators: Implications for Majorana Kramers Pair-Based Quantum Computing

Quantum spin Hall insulators (QSHIs) host helical electronic edge states that are protected from backscattering due to time-reversal symmetry (TRS). Despite considerable work investigating QSHI edge states, there is still an open question about their unexpected resilience to large magnetic fields where TRS is undoubtedly broken. In this work, we investigate the transport properties of helical edge states in a QSHI-superconductor (QSHI-SC) junction formed by a In⁢As(15 nm)/Ga⁢Sb(5 nm) double quantum well and a superconducting tantalum (Ta) constriction. We observe a robust conductance plateau up to 2 T, signaling resilient edge-state transport. Using a modified Landauer-Büttiker analysis, we find that the zero-field conductance is consistent with 98% Andreev-reflection probability owing to the high transparency of the (In⁢As/Ga⁢Sb)-Ta interface. Such resilience is consistent with the Dirac point for the edge states being buried in the bulk valence band. We further theoretically show that a buried Dirac point does not affect the robustness of the quasi-one-dimensional topological superconducting phase. We find that a buried Dirac point favors the hybridization of Majorana Kramers pairs (MKPs)—predicted to exist in a QSHI-SC constriction—and fermionic modes in the QSHI vacuum edge resulting in extended MKP states, highlighting the subtle role of buried Dirac points in probing MKPs.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

On the functional dependence of transition-potential coupled cluster

Orbital relaxation of the core region is a primary source of error in the computation of core ionization and core excitation energies. Recently, Transition-Potential Coupled Cluster (TP-CC) methods have been used to explicitly treat orbital relaxation using non-variational molecular orbitals determined by reoccupation of orbitals optimized for a fractional core occupation. The amount of fractional occupation is governed by parameter λ, and recommended values for accurate TP-CCSD and XTP-CCSD computations of carbon, nitrogen, oxygen, and fluorine K edges were previously determined. Herein, we explore the performance of several density functionals for generating the fractionally occupied orbitals used in TP-CCSD. These functionals include HF, BP86, BH&HLYP, B3LYP, M06-2X, and ωB97m-V. The fractionally occupied orbitals computed across the various functionals were subsequently employed as the initial orbitals for our TP-CCSD calculations of organic K-edge x-ray absorption and photoelectron spectra. Regardless of the functional used to generate the fractionally occupied orbitals, the TP-CCSD calculations yield accurate and comparable core ionization energies, core excitation energies, and oscillator strengths.

Coupled-cluster methods

Structural and electronic changes in L⁢i 2 ⁢Ru⁢O 3 induced by lithium intercalation

Despite extensive research on oxide battery cathodes that transcend classical cationic redox activity, the detailed interplay between structural transformations and electronic redox processes remains insufficiently understood. We report a detailed study of the sequential structural and electronic changes in Li 2 RuO 3 upon lithium intercalation, characterized by powder x-ray and neutron diffraction alongside Ru and O K-edge x-ray absorption spectroscopy (XAS), and guided by operando synchrotron x-ray diffraction. During delithiation, Li 2 RuO 3 evolves from a well-defined monoclinic state to a complex trigonal phase via multiple intermediate structures, marked by significant changes in Ru-O bond distances that closely track the transition from a classical cationic redox to an unconventional process centered at oxygen states. Armed with high-quality atomic structural descriptions, computational models of the O K-edge XAS closely reproduce the experimentally observed spectral shifts. Lastly, we relate observations of electrochemical hysteresis with concurrent changes in the pathways of structural and electronic transitions. In conclusion, our results not only clarify the mechanisms underpinning voltage hysteresis in a model for lattice oxygen redox but also underscore the importance of structural fidelity in modeling redox behavior when this type of complex reactivity is present.

Li, Haifeng [Univ. of Illinois, Chicago, IL (Unite

Elucidating the Discharge Behavior of Aqueous Zinc Sulfur Batteries in the Presence of Molybdenum(IV) Chalcogenide Catalyst: The Criticality of Interfacial Electrochemistry

The aqueous zinc-sulfur battery holds promise for significant capacity and energy density with low cost and safe operation based on environmentally benign materials. However, it suffers from the sluggish kinetics of the conversion reaction. Here, we highlight the efficacy of molybdenum(IV) sulfide (MoS 2 ) to reduce the overpotential of S-ZnS conversion in aqueous electrolytes and study the discharge products formed at the solid-solid and solid-liquid interfaces using experimental and theoretical approaches. Specifically, the MoS 2 -catalyzed electrochemical conversion reaction is characterized via ex situ X-ray diffraction (XRD), transmission electron microscopy (TEM) with energy dispersive spectroscopy (EDS), Raman spectroscopy, synchrotron-based Mo K-edge X-ray absorption spectroscopy (XAS), and in situ synchrotron-based X-ray computed tomography (XCT). Additionally, operando synchrotron-based S K-edge XAS and X-ray fluorescence (XRF) maps are collected to determine the spatial evolution of sulfur-based species at the electrode-electrolyte interface. Further, coupling the operando S K-edge XAS data with the simulated spectra and fitting the data suggested a possible ZnS 2 intermediate phase.

25 ENERGY STORAGE

PANDORA: A Parallel Dendrogram Construction Algorithm for Single Linkage Clustering on GPU

This paper introduces Pandora, a parallel algorithm for computing dendrograms, the hierarchical cluster trees for single linkage clustering (SLC). Current parallel approaches construct dendrograms by partitioning a minimum spanning tree and removing edges. However, they struggle with skewed, hard-to-parallelize real-world dendrograms. Consequently, computing dendrograms is the sequential bottleneck in HDBSCAN*[21], a popular SLC variant. Pandora uses recursive tree contraction to address this limitation. Pandora contracts nodes to construct progressively smaller trees. It computes the smallest contracted dendrogram and expands it by inserting contracted edges. This recursive strategy is highly parallel, skew-independent, work-optimal, and well-suited for GPUs and multicores. We develop a performance portable implementation of Pandora in Kokkos[31] and evaluate its performance on multicore CPUs and multi-vendor GPUs (e.g., Nvidia, AMD) for dendrogram construction in HDBSCAN*. Multithreaded Pandora is 2.2x faster than the current best-multithreaded implementation. Our GPU version achieves 6-20x speedup on AMD GPUs and 10-37x on NVIDIA GPUs over multithreaded Pandora. Pandora removes HDBSCAN*’s sequential bottleneck, greatly boosting efficiency, particularly with GPUs.

Sao, Piyush

Globus service enhancements for exascale applications and facilities

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications, particularly when leadership computing facilities are involved, can touch upon edge cases (e.g., terabyte files) that had not been a focus of previous Globus optimization work, which had emphasized rather the movement of many smaller (megabyte to gigabyte) files. We report here on how automated client-driven chunking can be used to accelerate both the movement of large files and the integrity checking operations that have proven to be essential for large data transfers. In conclusion, we present detailed performance studies that provide insights into the benefits of these modifications in a range of file transfer scenarios.

97 MATHEMATICS AND COMPUTING

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES

Liquid lithium divertor analysis using coupled plasma material interaction model

A liquid lithium divertor can improve performance of future fusion devices by creating efficient power exhaust and improving the energy confinement via pumping of the hydrogen isotopes. In addition, significantly higher heat fluxes can be handled if controlled vapor shielding is used to redistribute the divertor heat flux over a wider area. Design and optimization of such a system calls for an analysis model which includes a strong two-way coupling between the plasma and divertor material. The incoming plasma heat and particle flux will affect the divertor surface temperature, which is a defining factor of the lithium evaporative and sputtered flux going into the plasma. Results of the coupled model based on the plasma edge code SOLPS-ITER and the computational fluid dynamics (CFD) code ANSYS-CFX will be presented for different configurations. An analytical slab flow model is used as a heat transfer boundary condition for SOLPS, defining particle flux from the wall via calculation of the surface temperature. At the final step, results of the SOLPS analysis are verified using a 3D CFD magnetohydrodynamics (MHD) analysis which uses heat and particle flux from SOLPS as a boundary condition. In addition to plasma heat flux, both analytical and CFD temperature models include several plasma material interaction effects, such as lithium evaporation, condensation and sputtering based on deuterium target flux. New adatom sputtering model based on the available experimental data is presented. Analytical model is expanded to include free surface axisymmetric configurations. Results of parametric studies of the divertor configurations with different lithium inlet temperature and velocity will be presented leading to the optimal design resulting in the lowest possible lithium contamination in the core.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

The Heterogeneous Integration of Electronic Components

Heterogeneous integration (HI) of electronics components is broadly recognized as a powerful and crucial enabler for the continued growth of computing and communication. From 2010 onwards, the value of HI is increasingly visible in the advanced packaging used in artificial intelligence, high-performance computing, smartphones and communications product implementations. In this Perspective, we argue that HI is crucial to semiconductors and more broadly to the continued evolution of computing and communications. We use leading-edge advanced packaging examples to represent the value, advancements and opportunities for HI. To succeed, it is critical to develop comprehensive HI roadmaps that inform collaborations across the design, manufacturing and reliability spectrum between systems architects, packaging and semiconductor technologists to common goals. Although this article does not provide a full roadmap, we instead detail additional parameters for artificial intelligence, smartphone and other cellular communication devices, and their constituent building blocks including interconnects, power electronics, photonics, thermal management, reliability, modelling and co-design, to foster greater collaboration opportunities among academia, research laboratories and industry.

42 ENGINEERING

Cyclically symmetric radially self-similar phononic pseudocrystal isolator for broadband, ultrasonic vibration bandstop filtering

A 2D phononic pseudocrystal isolator exhibiting cyclic symmetry and radial self-similarity is measured and demonstrated to block a wide range of ultrasonic vibration. Measurements of longitudinal and shear wave blocking effects are made and compared with computational results. The use of the bandgap edge ratio is recommended for quantifying suppression in very-wide-bandgap materials. In conclusion, the upper-to-lower suppression edge frequency ratios of 3–4 are remarkably large for shear waves and even larger for longitudinal waves upper-to-lower suppression ratio (13 at 5 dB), such that 92.5% of frequencies in that range experience ≥ 5 dB of suppression.

Acoustic metamaterial

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data

Transforming Energy Through Computational Excellence: A View From NREL

At the National Renewable Energy Laboratory (NREL)—a U.S. Department of Energy laboratory—computational science, high-performance computing, applied mathematics, advanced computer science, visualization, and data play a pivotal role in advancing energy abundance, affordability, security, and reliability. From fundamental scientifc discovery to systems engineering and analysis, NREL researchers tackle market-relevant challenges to develop solutions for an independent energy system that is reliable, resilient and secure. Collaborative partnerships with industry, government, and academia ensure that our research remains cutting edge, impactful, applicable, and aligned with real-world energy needs. This special issue of Computing in Science & Engineering highlights exemplary NREL projects where computational tools and methodologies drive discovery and accelerate innovation in scalable and integrated energy systems. The featured articles explore the role of computational modeling, high-performance computing, generative AI, and adaptive computing in advancing independent energy solutions, optimizing sustainability research, and enhancing decision-making for energy solutions using a broad mix of energy technologies. Here, these contributions demonstrate how NREL’s computational research bridges the gap between theoretical advancements and practical implementation, emphasizing interdisciplinary collaboration and a commitment to innovation, with a focus on translating computational excellence into real-world impact, thus accelerate progress toward national energy goals. By showcasing cutting-edge research at the intersection of computational science and energy systems, this issue aims to inspire and inform researchers, practitioners, and policymakers dedicated to shaping a more reliable energy future.

97 MATHEMATICS AND COMPUTING