Search NASA⌕ Search

SEARCH · Search NASA

Results for “heterogeneous computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

CRCNS22 Learning Rules in the Hippocampus and their Mapping to Neuromorphic Systems (Final Technical Report)

Large scale biologically-realistic computational models are key to investigating the interplay between structure and function in nervous systems, thus paving the way to new clinical methods and neuro-inspired computing solutions. This project focuses on the hippocampus, in particular the CA3-CA1 regions, due to their role in associative learning and memory, pattern separation and completion, and spatial navigation. Investigations into the neuronal organization and learning rule(s) of this circuit can shed light into how declarative memories are formed, stored, recalled and forgotten and inform computational, experimental and clinical neuroscience work. Our project aims at developing a novel data-driven methodology supported by a broad heterogeneous base of neuroscience experimental knowledge and inspired from advances in computer science and engineering. Specifically, this work will benchmark existing and new learning rules within a full-scale spiking neural network simulation of the CA3-CA1 region. The model will be based on an open-source repository, called the Hippocampome, which contains neuronal morphologies, firing patterns, synapse probabilities, and most other required parameters for all known neuron types in the rodent hippocampal formation. The model will be first trained in a supervised fashion for associative memory tasks using backpropagation through time traditionally used in computer science, enhanced with a new technique called the surrogate gradient method. This optimization method will be used to obtain a global loss minimization, but it is not biologically inspired as it assumes the use of data not locally available to the synapses. However, we propose its use as a benchmarking tool, to compare the training performance of local biologically plausible and hardware-mappable learning rules at scale. New rules or combinations will be proposed and tested as needed, based on the obtained results. Progress in this area will also drive the development of novel hardware-mappable algorithms for continual lifelong learning and categorization of new events from few presented examples. This project goes beyond the existing state-of-the-art by looking at large scale realistic neuronal circuits as networks trainable via global optimization methods such as surrogate gradient descent. The objective function of the brain that supports learning is largely unknown, but it is likely that it operates through local learning rules. Studying network trajectories around local minima as proposed in this work represents a useful strategy for understanding whether a network is training by using a specific (set of) learning rule(s). Starting from a completely untrained network is a challenging test since it is difficult to determine how the learning rule affects the trajectory of the network. This interdisciplinary project will help understand what rule governs learning in these regions or if multiple learning rules are involved. The work will develop a robust methodology to measure if the network is converging to the target solution, oscillating around it, or diverging away.

59 BASIC BIOLOGICAL SCIENCES↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Computational prediction of dielectric breakdown strength of a transformer paper in oil with uncertainty quantification

The determination of the dielectric breakdown strengths of microstructurally heterogeneous materials has been a primarily experimental endeavor. We report the development of a microstructure-level model for computationally predicting the breakdown strength and analyzing the interactions between electromagnetic pulses (EMP) and the constituents in a composite of cellulose-based paper and mineral oil found in electrical transformers. The model allows explicit simulation of the material breakdown process by tracking the transition of dielectric constituents from non-conductive to conductive states. The focus is on the electric fields induced in the materials and the overall conditions for dielectric breakdown (defined as the onset of avalanche) caused by the electric field induced in the composite. Responses to three distinct pulse shapes, i.e., Steep Front (SF), Lightning (L), and AC with spectra spanning 60–9 × 105 Hz are considered. It is found that the breakdown strength of the material is significantly affected by microstructure heterogeneities, the spatial variations of the constituent properties, and the pulse shapes. A probabilistic characterization of the breakdown strength is computationally obtained and compared with experimental measurements. Although one particular material is analyzed, the model and approach are applicable to other heterogeneous materials as well.

breakdowns↗

Development of a River Dynamical Core for E3SM to simulate compound flooding on Exascale-class heterogeneous supercomputers

Flooding events pose significant risk to human life, property, and infrastructure. Physically-consistent quantification of altered flood risks in global models requires hyper-resolution (~1 km) or fine flood simulations using two-dimensional (2D) physics schemes, both of which are unavailable in the current generation Earth System Models. Here, in this work, we have developed the River Dynamical Core (RDycore), which is an open-source, 2D shallow water equation (SWE) library for the U.S. Department of Energy's Energy Exascale Earth System Model (E3SM). RDycore uses PETSc and libCEED libraries that allows it to run efficiently on CPUs and GPUs, as well as select a time-integration algorithm at runtime without requiring any code modifications. RDycore achieves spatial error convergence rates for problems with analytical and manufactured solutions similar to those reported previously in the literature, or consistent with the implemented first-order spatial discretization scheme. RDycore's accuracy in predicting flooding for a well-studied dam break problem is comparable to existing SWE models. For a problem with 471 million grid cells, RDycore achieves a speedup of 6.6x and 7.6x on GPUs compared to CPUs when using 320 compute nodes on DOE's Perlmutter and Frontier supercomputers, respectively. The one-way coupling of the RDycore library within E3SM is demonstrated by performing multiple 5-day flooding simulations during Hurricane Harvey driven by five precipitation datasets. The E3SM--RDycore simulations at 30 m spatial resolution accurately simulate maximum water height during the hurricane when benchmarked against a previously published study and achieve a speedup of 15x (Perlmutter) and 21x (Frontier) on GPUs relative to CPUs. The work presented here is the foundational step in providing hardware and algorithmic portability framework for simulating kilometer-scale river dynamics within E3SM.

Flood Simulation↗

Packaging HEP Heterogeneous Mini-apps for Portable Benchmarking and Facility Evaluation on Modern HPCs

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to make informed decisions as to where they wish to devote resources. In its first phase, the High Energy Physics Center for Computational Excellence (HEP-CCE) produced portable versions of a number of heterogeneous HEP mini-apps, such as p2r, FastCaloSim, Patatrack and the WireCell Toolkit, that exercise a broad range of GPU characteristics, enabling cross platform and facility benchmarking and evaluation. However, these miniapps still require a significant amount of manual intervention to deploy on a new facility. We present our work in developing turn-key deployments of these mini-apps, where by means of containerization and automated configuration and build techniques such as Spack, we are able to quickly test new hardware, software, environments and entire facilities with minimal user intervention, and then track performance metrics over time.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Packaging HEP Heterogeneous Mini-apps for Portable Benchmarking and Facility Evaluation on Modern HPCs

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to make informed decisions as to where they wish to devote resources. In its first phase, the High Energy Physics Center for Computational Excellence (HEP-CCE) produced portable versions of a number of heterogeneous HEP mini-apps, such as \ptor, FastCaloSim, Patatrack and the WireCell Toolkit, that exercise a broad range of GPU characteristics, enabling cross platform and facility benchmarking and evaluation. However, these mini-apps still require a significant amount of manual intervention to deploy on a new facility. We present our work in developing turn-key deployments of these mini-apps, where by means of containerization and automated configuration and build techniques such as Spack, we are able to quickly test new hardware, software, environments and entire facilities with minimal user intervention, and then track performance metrics over time.

Atif, Mohammad [Brookhaven] (ORCID:000000026889770↗

Segmentation of RDX and TNT in X‐Ray Computed Tomography Reconstructions of Melt‐Cast Explosives

ABSTRACT Three‐dimensional mesoscale characterization of heterogeneous melt‐cast high explosives is challenging because of the difficulty differentiating binder from explosive crystals: two functionally different materials which are typically similar in density by design. Here, we report an algorithm which can differentiate hexahydro‐1,3,5‐trinitro‐1,3,5‐triazine (RDX) from 2,4,6‐trinitrotoluene (TNT) in x‐ray computed tomography (CT) volumes with tens of microns resolution. This method allows us to quantify RDX/TNT content, porosity, and RDX domain size. We calibrated the segmentation algorithm using simulated x‐ray CT volumes containing object models of RDX crystals within a TNT matrix. We then segmented and analyzed CT data for Composition B (Comp B), a 60/40 RDX/TNT mixture, and Cyclotol, a 75/25 RDX/TNT mixture. We examined melt‐cast samples fabricated with 100% theoretical maximum density (TMD) and 85% TMD. For the 100% TMD Comp B and Cyclotol samples, the RDX content values calculated by segmentation were 3% and 9% lower, respectively, than the values measured by high‐performance liquid chromatography on material from the same synthesis lots. This result is consistent with the expected underreporting of RDX content resulting from x‐ray CT resolution limits on RDX particles with diameters smaller than 25 µm. The 85% TMD samples were less accurately segmented with our algorithm due to the confounding presence of voids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Control of Energy Transport and Transduction in Photosynthetic Down-Conversion (Final Technical Report)

The research conducted under this contract focused on three areas related to control of energy transport and transduction in natural and synthetic light harvesting systems: (1) experimental and theoretical studies exploring the use of ultrafast laser pulse shaping to manipulate the initially excited states of multi-chromophore biological light harvesting networks and influence the early time energy transfer and electronic and vibrational relaxation pathways, (2) fundamental theoretical developments aimed at computing and analyzing the dynamics of excitations of complex heterogeneous environments and nano structures, and (3) joint experimental and theoretical design studies of artificial light harvesting materials exploring the influence of incorporating different types of organic layers in low dimensional lead halide perovskite materials on their hot carrier relaxation dynamics, and theoretical studies of the design of nano structured light harvesting antenna systems. These research projects supported the training of three theoretical and computational graduate students, one experimental graduate student, and one experimental post-doctoral researcher. The research has been published in seven peer reviewed papers and presented at several international meetings and workshops.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BoBa

BoBa is a C++ software library for working with large matrices, tensors, and tensor decompositions. The library provides tools for dense matrix and tensor operations, tensor decompositions, and tensor decomposition methods that support modern CPU and GPU architectures. It includes portable abstractions for linear algebra, tensor algebra, and multidimensional computation. BoBa is intended for scientific computing applications that involve large multidimensional data sets or high dimensional mathematical models. Its capabilities support tasks such as data compression, linear algebra, efficient numerical computation, and the development of scalable algorithms for heterogeneous hardware. Tutorials, tests, and example applications are included to help users learn and apply the library.

Yao, Jin [Lawrence Livermore National Laboratory (↗

Extending SEER for Extreme Heterogeneity

Heterogeneous and multi-device nodes are increasingly common in high-performance computing and data centers, yet existing programming models often lack simple, transparent, and portable support for these diverse architectures. The main contribution of this work is the development of novel SEER capabilities to address this challenge by providing a descriptive programming model that allows applications to seamlessly leverage heterogeneous nodes across various device types. SEER uses efficient memory management and can select the proper device[s] depending on the computational cost of the applications. This is completely transparent to the programmer, thereby providing a highly productive programming environment. Integrating extreme heterogeneity into the SEER library as shown with the use of NVIDIA and AMD GPUs simultaneously allows it to expand and exploit the performance possibilities. Our analysis based on the well-known Conjugate Gradient algorithm reports accelerations above 1.5 × on computationally demanding steps of such an algorithm by using both architectures simultaneously.

Teranishi, Keita [ORNL] (ORCID:0000000166472690)↗

A review of thermo-hydro-mechanical modeling of coupled processes in fractured rock: From continuum to discontinuum perspective

Coupled thermo-hydro-mechanical (THM) processes in fractured rock are playing a crucial role in geoscience and geoengineering applications. Diverse and conceptually distinct approaches have emerged over the past decades in both continuum and discontinuum perspectives leading to significant progress in their comprehending and modeling. This review paper offers an integrated perspective on existing modeling methodologies providing guidance for model selection based on the initial and boundary conditions. By comparing various models, one can better assess the uncertainties in predictions, particularly those related to the conceptual models. The review explores how these methodologies have significantly enhanced the fundamental understanding of how fractures respond to fluid injection and production, and improved predictive capabilities pertaining to coupled processes within fractured systems. It emphasizes the importance of utilizing advanced computational technologies and thoroughly considering fundamental theories and principles established through past experimental evidence and practical experience. The selection and calibration of model parameters should be based on typical ranges and applied to the specific conditions of applications. The challenges arising from inherent heterogeneity and uncertainties, nonlinear THM coupled processes, scale dependence, and computational limitations in representing field scale fractures are discussed. Realizing potential advances on computational capacity calls for methodical conceptualization, mathematical modeling, selection of numerical solution strategies, implementation, and calibration to foster simulation outcomes that intricately reflect the nuanced complexities of geological phenomena. Future research efforts should focus on innovative approaches to tackle the hurdles and advance the state-of-the-art in this critical field of study.

Coupling scheme↗

Bridging molecular-scale interfacial science with continuum-scale models

Solid–water interfaces are crucial for clean water, conventional and renewable energy, and effective nuclear waste management. However, reflecting the complexity of reactive interfaces in continuum-scale models is a challenge, leading to oversimplified representations that often fail to predict real-world behavior. This is because these models use fixed parameters derived by averaging across a wide physicochemical range observed at the molecular scale. Recent studies have revealed the stochastic nature of molecular-level surface sites that define a variety of reaction mechanisms, rates, and products even across a single surface. To bridge the molecular knowledge and predictive continuum-scale models, we propose to represent surface properties with probability distributions rather than with discrete constant values derived by averaging across a heterogeneous surface. This conceptual shift in continuum-scale modeling requires exponentially rising computational power. By incorporating our molecular-scale understanding of solid–water interfaces into continuum-scale models we can pave the way for next generation critical technologies and novel environmental solutions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Adaptive Interface-PINNs (AdaI-PINNs) for transient diffusion: Applications to forward and inverse problems in heterogeneous media

We model transient diffusion in heterogeneous materials using a novel physics-informed neural networks framework (PINNs) termed Adaptive interface physics-informed neural networks or AdaI-PINNs (Roy et al. arXiv preprint arXiv:2406.04626, 2024). AdaI-PINNs utilize different activation functions with trainable slopes tailored to each material region within the computational domain, allowing for a fully automated and adaptive PINNs approach to model interface problems with strongly and weakly discontinuous solutions. To enhance its performance in highly heterogeneous transient diffusion systems, we prescribe a suite of robust practices, including appropriate non-dimensionalization of equations, a biased sampling method, Glorot initialization, and the hard enforcement of boundary and initial conditions. Here we evaluate the efficacy of the proposed method on several benchmark forward and inverse problems. Comparative studies on one-dimensional and two-dimensional benchmark problems reveal that the modified AdaI-PINNs outperform its unmodified counterpart, achieving root-mean-square errors that are at least two orders of magnitude better in forward problems. For inverse problems, the maximum errors in the approximated diffusion coefficients by modified AdaI-PINNs are four orders of magnitude better than those of the unmodified version. Additionally, modified AdaI-PINNs demonstrate improved stability in problems with large material mismatches.

42 ENGINEERING↗

Posterior Covariance Matrix Approximations

Here, the Davis equation of state (EOS) is commonly used to model thermodynamic relationships for high explosive (HE) reactants. Typically, the parameters in the EOS are calibrated, with uncertainty, using a Bayesian framework and Markov Chain Monte Carlo (MCMC) methods. However, MCMC methods are computationally expensive, especially for complex models with many parameters. This paper provides a comparison between MCMC and less computationally expensive Variational methods (Variational Bayesian and Hessian Variational Bayesian) for computing the posterior distribution and approximating the posterior covariance matrix based on heterogeneous experimental data. All three methods recover similar posterior distributions and posterior covariance matrices. This study demonstrates that for this EOS parameter calibration application, the assumptions made in the two Variational methods significantly reduce the computational cost but do not substantially change the results compared to MCMC.

97 MATHEMATICS AND COMPUTING↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Evaluating crown scorch predictions from a computational fluid dynamics wildland fire simulator

Abstract Background Crown scorch—the heating of live leaves, needles, and buds in the vegetative canopy to lethal temperatures without widespread combustion—is one of the most common fire effects shaping post-fire canopies. Despite the ability of computational fluid dynamic models to finely resolve fire activity and buoyant plume dynamics including heterogenous 3D distributions of forest canopy heating, these models have had only limited use in simulating fire effects and have not been used to evaluate crown scorch. Here, we demonstrate a method of evaluating crown scorch using a computational fluid dynamics model, FIRETEC, and validate this approach by simulating the experiments that were used to develop Van Wagner’s 1973 crown scorch model. Results The average scorch height prediction from FIRETEC compares well with the empirical model derived by Van Wagner, which is the most widely used empirical model for crown scorch. We further find that the 3D buoyant plume dynamics from a steady and homogeneous idealized heat source on the ground results in a spatially heterogenous crown scorch pattern reflecting complex heating dynamics that are best represented by percent scorch rather than height of scorch. Conclusions The ability of the computational fluid dynamics model to capture variation in crown scorch due to 3D buoyant plume dynamics provides direct links between forest structure, fire behavior, and fire effects that can be used by forest managers and researchers to better understand how fires result in crown damage under various environmental and management scenarios.

54 ENVIRONMENTAL SCIENCES↗

Learning a general model of single phase flow in complex 3D porous media

Modeling effective transport properties of 3D porous media, such as permeability, at multiple scales is challenging as a result of the combined complexity of the pore structures and fluid physics—in particular, confinement effects which vary across the nanoscale to the microscale. While numerical simulation is possible, the computational cost is prohibitive for realistic domains, which are large and complex. Although machine learning (ML) models have been proposed to circumvent simulation, none so far has simultaneously accounted for heterogeneous 3D structures, fluid confinement effects, and multiple simulation resolutions. By utilizing numerous computer science techniques to improve the scalability of training, we have for the first time developed a general flow model that accounts for the pore-structure and corresponding physical phenomena at scales from Angstrom to the micrometer. Using synthetic computational domains for training, our ML model exhibits strong performance (R 2 = 0.9) when tested on extremely diverse real domains at multiple scales.

36 MATERIALS SCIENCE↗

Learning a General Model of Single Phase Flow in Complex 3D Porous Media

Modeling effective transport properties of 3D porous media, such as permeability, at multiple scales is challenging as a result of the combined complexity of the pore structures and fluid physics—in particular, confinement effects which vary across the nanoscale to the microscale. While numerical simulation is possible, the computational cost is prohibitive for realistic domains, which are large and complex. Although machine learning (ML) models have been proposed to circumvent simulation, none so far has simultaneously accounted for heterogeneous 3D structures, fluid confinement effects, and multiple simulation resolutions. By utilizing numerous computer science techniques to improve the scalability of training, we have for the first time developed a general flow model that accounts for the pore-structure and corresponding physical phenomena at scales from Angstrom to the micrometer. Using synthetic computational domains for training, our ML model exhibits strong performance (R 2 = 0.9) when tested on extremely diverse real domains at multiple scales.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗