Search NASA⌕ Search

SEARCH · Search NASA

Results for “kernel design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Physics-constrained Gaussian process model for prediction of hydrodynamic interactions between wave energy converters in an array

To improve the efficiency of wave farms and achieve maximum power generation, the layout of wave energy converters (WECs) in an array needs to be carefully designed so that the hydrodynamic interactions can be positively exploited. For this, the hydrodynamic characteristics of the WEC array in different layouts need to be calculated. However, such calculations using numerical models usually entail significant computational cost, especially for large arrays of WECs. To address the computational challenge, a physics-constrained Gaussian process (GP) model is proposed to replace the original expensive numerical model and predict the hydrodynamic characteristics of the WECs for any array layout. By exploring the relationship between the WEC array (i.e., the input) and different hydrodynamic characteristics (i.e., the output), here we summarize a set of physical constraints/features, including invariance, symmetry, and additivity. This prior knowledge about the input-output relationship is then directly embedded in the constructed GP model through the design of physics-constrained kernels. In particular, a double-sum invariant kernel is first developed to incorporate the invariance and symmetry features, and then an additive kernel is developed to incorporate the additive feature of the problem. The invariant kernel and the additive kernel are then integrated to construct the physics-constrained GP model. Compared to the standard GP model, the proposed physics-constrained GP models require less training data to achieve the desired accuracy in predicting the hydrodynamic characteristics and are also less vulnerable to the curse of dimensionality (i.e., good scalability for large arrays) due to the use of an additive kernel. The efficiency, accuracy, and scalability of the proposed approach are demonstrated through an application to predict the hydrodynamic characteristics for WEC arrays of different sizes and layouts.

16 TIDAL AND WAVE POWER↗

Exact Gaussian processes for massive datasets via non-stationary sparsity-discovering kernels

Abstract A Gaussian Process (GP) is a prominent mathematical framework for stochastic function approximation in science and engineering applications. Its success is largely attributed to the GP’s analytical tractability, robustness, and natural inclusion of uncertainty quantification. Unfortunately, the use of exact GPs is prohibitively expensive for large datasets due to their unfavorable numerical complexity of $$O(N^3)$$ O ( N 3 ) in computation and $$O(N^2)$$ O ( N 2 ) in storage. All existing methods addressing this issue utilize some form of approximation—usually considering subsets of the full dataset or finding representative pseudo-points that render the covariance matrix well-structured and sparse. These approximate methods can lead to inaccuracies in function approximations and often limit the user’s flexibility in designing expressive kernels. Instead of inducing sparsity via data-point geometry and structure, we propose to take advantage of naturally-occurring sparsity by allowing the kernel to discover—instead of induce—sparse structure. The premise of this paper is that the data sets and physical processes modeled by GPs often exhibit natural or implicit sparsities, but commonly-used kernels do not allow us to exploit such sparsity. The core concept of exact, and at the same time sparse GPs relies on kernel definitions that provide enough flexibility to learn and encode not only non-zero but also zero covariances. This principle of ultra-flexible, compactly-supported, and non-stationary kernels, combined with HPC and constrained optimization, lets us scale exact GPs well beyond 5 million data points.

97 MATHEMATICS AND COMPUTING↗

Deep energy-pressure regression for a thermodynamically consistent EOS model

Abstract In this paper, we aim to explore novel machine learning (ML) techniques to facilitate and accelerate the construction of universal equation-Of-State (EOS) models with a high accuracy while ensuring important thermodynamic consistency. When applying ML to fit a universal EOS model, there are two key requirements: (1) a high prediction accuracy to ensure precise estimation of relevant physics properties and (2) physical interpretability to support important physics-related downstream applications. We first identify a set of fundamental challenges from the accuracy perspective, including an extremely wide range of input/output space and highly sparse training data. We demonstrate that while a neural network (NN) model may fit the EOS data well, the black-box nature makes it difficult to provide physically interpretable results, leading to weak accountability of prediction results outside the training range and lack of guarantee to meet important thermodynamic consistency constraints. To this end, we propose a principled deep regression model that can be trained following a meta-learning style to predict the desired quantities with a high accuracy using scarce training data. We further introduce a uniquely designed kernel-based regularizer for accurate uncertainty quantification. An ensemble technique is leveraged to battle model overfitting with improved prediction stability. Auto-differentiation is conducted to verify that necessary thermodynamic consistency conditions are maintained. Our evaluation results show an excellent fit of the EOS table and the predicted values are ready to use for important physics-related tasks.

97 MATHEMATICS AND COMPUTING↗

Mathematical nuances of Gaussian process-driven autonomous experimentation

Abstract The fields of machine learning (ML) and artificial intelligence (AI) have transformed almost every aspect of science and engineering. The excitement for AI/ML methods is in large part due to their perceived novelty, as compared to traditional methods of statistics, computation, and applied mathematics. But clearly, all methods in ML have their foundations in mathematical theories, such as function approximation, uncertainty quantification, and function optimization. Autonomous experimentation is no exception; it is often formulated as a chain of off-the-shelf tools, organized in a closed loop, without emphasis on the intricacies of each algorithm involved. The uncomfortable truth is that the success of any ML endeavor, and this includes autonomous experimentation, strongly depends on the sophistication of the underlying mathematical methods and software that have to allow for enough flexibility to consider functions that are in agreement with particular physical theories. We have observed that standard off-the-shelf tools, used by many in the applied ML community, often hide the underlying complexities and therefore perform poorly. In this paper, we want to give a perspective on the intricate connections between mathematics and ML, with a focus on Gaussian process-driven autonomous experimentation. Although the Gaussian process is a powerful mathematical concept, it has to be implemented and customized correctly for optimal performance. We present several simple toy problems to explore these nuances and highlight the importance of mathematical and statistical rigor in autonomous experimentation and ML. One key takeaway is that ML is not, as many had hoped, a set of agnostic plug-and-play solvers for everyday scientific problems, but instead needs expertise and mastery to be applied successfully. Graphical abstract

97 MATHEMATICS AND COMPUTING↗

Adrastea: An Efficient FPGA Design Environment for Heterogeneous Scientific Computing and Machine Learning

We present Adrastea, an efficient FPGA design environment for developing scientific machine learning applications. FPGA development is challenging, from deployment, proper toolchain setup, programming methods, interfacing FPGA kernels, and more importantly, the need to explore design space choices to get the best performance and area usage from the FPGA kernel design. Adrastea provides an automated and scalable design flow to parameterize, implement, and optimize complex FPGA kernels and associated interfaces. We show how virtualization of the development environment via virtual machines is leveraged to simplify the setup of the FPGA toolchain while deploying the FPGA boards and while scaling up the automated design space exploration to leverage multiple machines concurrently. Adrastea provides an automated build and test environment of FPGA kernels. By exposing design space hyper-parameters, Adrastea can automatically search the design space in parallel to optimize the FPGA design for a given metric, usually performance or area. Adrastea simplifies the task of interfacing with the FPGA kernels with a simplified interface API. To demonstrate the capabilities of Adrastea, we implement a complex random forest machine learning kernel with 10,000 input features while achieving extremely low computing latency without loss of prediction accuracy, which is required by a scientific edge application at SNS. We also demonstrate Adrastea using an FFT kernel and show that for both applications Adrastea is able to systematically and efficiently evaluate different design options, which reduced the time and effort required to develop the kernel from months of manual work to days of automatic builds.

Young, Aaron↗

Kernel Buffer Volume Fraction Margin of the AGR Designed Fuel Particle

Modeling results used to assess the fuel performance of the TRISO-coated fuel particles as a function of kernel/buffer volume fraction include SiC tangential stress, formation of the buffer/IPyC gap, particle temperature profile, internal particle pressure, fission gas released from the kernel, probability of fuel particle failure, and fission product diffusion. These results were evaluated at two burnup levels and irradiation temperatures to bound expected steady-state irradiation conditions. In general, increasing the kernel/buffer volume fraction increases the SiC stress and subsequently the failure probability of a fuel particle when compared to the AGR designed particle. There was little impact on the fission product diffusion through the particle as the kernel/buffer volume fraction increased.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Working with Bézier Curves as bases for Functional Expansion Tallies

Functional expansion tallies (FETs) are powerful tools for getting more information per history from Monte Carlo simulations, but in the past they have been constrained to orthogonal bases. Bézier curves are used widely in computer aided design (CAD) geometry kernels and could be well-suited for FETs due to their ability to assume many arbitrary shapes, but they use nonorthogonal bases. Recent developments in 2021 have made nonorthogonal FETs possible. The convergence of Bézier curve FETs in both polynomial order and number of samples is explored in this work. It is shown that these bases are well-suited for representing normal distributions and this opens the door to the possibility of other CAD-derived FET bases.

97 - MATHEMATICS AND COMPUTING↗

Hierarchical Speed Planner for Automated Vehicles: A Framework for Lagrangian Variable Speed Limit in Mixed-Autonomy Traffic

Here, this article presents a novel hierarchical speed planning framework for variable speed limits in mixed-autonomy traffic environments, leveraging server-side macroscopic control and vehicle-side microscopic execution. The framework integrates real-time traffic state estimation (TSE) and reinforcement learning (RL)-based control to mitigate congestion and improve traffic flow. A TSE enhancement module combines macroscopic data from sources like INRIX with high-resolution observations from connected autonomous vehicles (CAVs), enabling predictive modeling to address latency and noise. The target speed design module employs kernel smoothing and a buffer zone strategy to optimize traffic density and flow around bottlenecks. The proposed system was validated in the largest open-road test to date with 100 CAVs, demonstrating an overall 8% traffic density decrease, with a specific decrease of 7% upstream, 10% downstream, and a 52% decrease during the congestion formation phase at bottlenecks.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Physics-Informed Gaussian Process Inference of Liquid Structure from Scattering Data

We present a nonparametric Bayesian framework to infer radial distribution functions from experimental scattering measurements with uncertainty quantification using nonstationary Gaussian processes. The Gaussian process prior mean and kernel functions are designed to mitigate well-known numerical challenges with the Fourier transform, including discrete measurement binning and detector windowing, while encoding fundamental yet minimal physical knowledge of the liquid structure. We demonstrate uncertainty propagation of the Gaussian process posterior to unmeasured quantities of interest. Experimental radial distribution functions of liquid argon and water with uncertainty quantification are provided as both a proof of principle for the method and a benchmark for molecular models.

Chemical structure↗

Architecture-Aware Models of AI Engines for High-Performance Matrix Matrix Multiplication

The AI Engine (AIE) architecture, available in systems from mobile SoCs to server-class FPGAs, aims to efficiently execute AI/ML tasks through a two-dimensional array of compute tiles. Previous work on AIEs has explored different approaches to mapping computation across spatial arrays, but the compute kernel running on each tile has not been the focus. Additionally, the AIE-ML architecture introduces memory tiles and omits programmable logic, requiring new approaches to staging and moving data throughout the array. In this work we update analytical models developed for CPUs to produce the design of high performance kernels while introducing new model considerations such as memory structure, throughput, and latency as required by the AIE hardware. We evaluate our models by developing AIE-ML kernels for matrix multiplication in low-precision data types showing performance up to 95% of compute peak for the kernel when data resides in local memory and above 90% of compute peak when data resides in main memory.

Binder, Elliott D. [Carnegie Mellon University, Pi↗

HetArch: Heterogeneous Microarchitectures for Superconducting Quantum Systems

Noisy Intermediate-Scale Quantum Computing (NISQ) has dominated headlines in recent years, with the longer-term vision of Fault-Tolerant Quantum Computation (FTQC) offering significant potential but at currently intractable resource costs and quantum error correction (QEC) overheads. For problems of interest, FTQC will require millions of physical qubits with long coherence times, high-fidelity gates, and compact sizes to surpass classical systems. Just as heterogeneous specialization has offered scaling benefits in classical computing, it is likewise gaining interest in FTQC. However, systematic use of heterogeneity in either hardware or software elements of FTQC systems remains a serious challenge due to the vast design space and the variable physical constraints. This paper meets the challenge of making heterogeneous FTQC design practical by introducing HetArch, a toolbox for designing heterogeneous quantum systems, and using it to explore heterogeneous design scenarios. Using a hierarchical approach, we successively break quantum algorithms into smaller operations (akin to classical application kernels), thus greatly simplifying the design space and resulting tradeoffs. Specializing to superconducting systems, we then design optimized heterogeneous hardware composed of varied superconducting devices, abstracting physical constraints into design rules that enable devices to be assembled into standard cells optimized for specific operations, which, in turn, form heterogeneous modules optimized for quantum subroutines. Finally, we provide a heterogeneous design space exploration framework which reduces the simulation burden by a factor of 10^4 or more and allows us to characterize optimal design points. We use these techniques to design superconducting quantum modules for entanglement distillation, error correction, and code teleportation, reducing error rates by 2.6×, 10.7×, and 3.4× compared to homogeneous systems.

Quantum Computing, Quantum Physics, Computer Archi↗

TwoFold: Highly accurate structure and affinity prediction for protein-ligand complexes from sequences

We describe our development of ab initio protein-ligand binding pose prediction models based on transformers and binding affinity prediction models based on the neural tangent kernel (NTK). Folding both protein and ligand, the TwoFold models achieve efficient and quality predictions matching state-of-the-art implementations while additionally reconstructing protein structures. In conclusion, solving NTK models points to a new use case for highly optimized linear solver benchmarking codes on HPC.

60 APPLIED LIFE SCIENCES↗

Batched Sparse Linear Algebra (Final Report for Subcontract B648960)

This report finalizes design specifications for developing batched kernels for small tensor operations for unassembled matrix-free iterative solvers, batched solvers for partially assembled operators, and batched solvers with support for various sparse formats. The outcome of the project milestones is a set of interfaces to Batched Sparse LA solvers running on hardware accelerators for use in ECP Libraries and Applications. It is part of the development of sparse batched kernels, solvers/preconditioners as well as creating interoperability in xSDK libraries with sparse and dense batched functions to benefit ECP applications. The participants included representatives from ECP libraries (not limited to the xSDK project), applications, and vendors (AMD, Intel, and NVIDIA). Batched sparse linear algebra solvers form the new frontier for algorithmic development and performance engineering. Many applications (ECP and non-ECP alike) require simultaneous solutions of small linear systems of equations that are structurally sparse. To move towards high hardware utilization, it is important to provide these applications with appropriate interfaces to efficient batched sparse solvers running on modern hardware accelerators. We present interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the software portable between the major hardware accelerators from AMD, Intel, and NVIDIA. The presented interface specifications includes batched band, sparse iterative, and sparse direct solvers. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, SUNDIALS, and SuperLU_dist.

97 MATHEMATICS AND COMPUTING↗

Assimilating partial observation to enhance feedback control of stochastic dynamical systems

Here, in this paper, we present a novel methodology to tackle feedback optimal control problems in scenarios where the exact state of the controlled process is unknown. It integrates data assimilation techniques and optimal control solvers to manage partial observation of the state process, a common occurrence in practical scenarios. Traditional stochastic optimal control methods assume full state observation, which is often not feasible in real-world fluid dynamics control problems. Our approach underscores the significance of utilizing observational data to inform control policy design. Specifically, we introduce a kernel learning backward stochastic differential equation (SDE) filter to enhance data assimilation efficiency and propose a sample-wise stochastic optimization method within the stochastic maximum principle framework. We demonstrate the efficacy and accuracy of our method in the control of advection-diffusion-reaction flow problem and the Dubins airplane maneuvering problem with model uncertainty.

data driven↗

Micromechanical Properties of the SiC and Pyrolytic Carbon Layers in Tristructural-Isotropic Coated Particles

Tristructural isotropic (TRISO) coated particle fuel was initially developed for high temperature gas cooled reactors (HTGR) and has been proposed for several other advanced reactor concepts. The design of TRISO particles focuses on preventing the release of fission products in normal and off-normal reactor conditions. The particle design features an actinide bearing fuel kernel that is surrounded by three pyrolytic carbon (PyC) layers and a silicon carbide layer (SiC). The mechanical stability of the particle and fission product retention for both metallic and gaseous fission products depend on the SiC layer. Post irradiation examination (PIE) of TRISO fuel from the first two US DOE Advanced Gas Reactor Fuel Development and Qualification Program irradiation campaigns, AGR-1 and AGR-2, had identified a low rate of particles exhibiting cracking in the SiC layer that did not propagate across the SiC layer. While cracking in the SiC is rare for test conditions and particles associated with the AGR program, understanding the stress state and mechanical properties of the SiC and PyC layers related to particle architecture can aid predicting thermomechanical response of TRISO fuel under the prescribed operation envelope and beyond as well as aiding in the development of similar fuel concepts for other advanced reactors. The presented investigation shows the relationship of the mechanical properties and mechanical response (e.g., understanding crack propagation) of the SiC and PyC layers relative to position within the particle. Testing was conducted on the inner and outer PyC layers of TRISO particles to quantify differences in mechanical behavior.

Montoya, Katherine [ORNL] (ORCID:0000000326955086)↗

Boosting RDataFrame performance with transparent bulk event processing

RDataFrame is ROOT’s high-level interface for Python and C++ data analysis. Since it first became available, RDataFrame adoption has grown steadily and it is now poised to be a major component of analysis software pipelines for LHC Run 3 and beyond. Thanks to its design inspired by declarative programming principles, RDataFrame enables the development of highperformance, highly parallel analyses without requiring expert knowledge of multi-threading and I/O: user logic is expressed in terms of self-contained, small computation kernels tied together by a high-level API. This design completely decouples analysis logic from its actual execution, and opens several interesting avenues for workflow optimization. In particular, in this work we explore the benefits of moving internal data processing from an event-by-event to a bulkby-bulk loop. This refactoring dramatically reduces the framework’s runtime overheads; in collaboration with the I/O layer it improves data access patterns; it exposes information that optimizing compilers might use to auto-vectorize the invocation of user-defined computations; finally, while existing user-facing interfaces remain unaffected, it becomes possible to additionally offer interfaces that explicitly expose bulks of events, useful e.g. for the injection of GPU kernels into the analysis workflow. In order to inform similar future R&D, design challenges will be presented, as well as an investigation of the relevant timememory trade-off backed by novel performance benchmarks.

Guiraud, Enrico↗

Multi-scale fission product release model with comparison to AGR data

TRistructural ISOtropic (TRISO) particle fuel is central to several advanced, high-temperature reactor designs. Each particle consists of a fuel kernel encapsulated by three layers of carbon and ceramics that prevent the release of fission products and ensure physical integrity. Despite outstanding retention properties, fission product release has been observed from intact particles. To better understand and quantify fission product release from TRISO particles, a multiscale, mechanistic model of fission product transport is being developed by the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program. Previous work focused on silver (Ag) transport and improved Ag release predictions. The work described in this report builds on this experience to better understand cesium (Cs) transport in silicon carbide (SiC), the main barrier to the release of fission products. Atomistic simulations provide bulk and grain boundary (GB) Cs diffusivities in SiC, which are used by phase field simulations in the mesoscale code Marmot to determine the temperature, microstructure, and irradiation-dependent Cs diffusivity at the mesoscale in SiC. This approach attributes the different temperature regimes experimentally observed for Cs diffusivities in SiC to a transition from bulk-dominated diffusivity at high temperatures to a GB-dominated regime at low temperatures, providing new insight. The multiscale, mechanistic effective diffusivity is then implemented in the fuel performance code BISON and further validated by comparing Cs release predictions from Advanced Gas Reactor (AGR)-1 and AGR-2 post-irradiation measurements. The new model improves BISON’s predictability. This document also reports improvements made on Ag transport modeling by accounting for different GB types having different diffusivities. Moreover, this report details preliminary efforts to model palladium (Pd) attack of the SiC at the mesoscale using a phase field approach. Pd attack and its impact on accelerated Ag transport remains a misunderstood phenomenon, and we use the model to demonstrate that the formation of lamellae that has been observed in experiments can be explained by the reaction of Pd with SiC to form alternating layers of graphite and Pd 2 Si. This effort aims to improve our understanding of the reaction and eventually provide a model for BISON to account for Pd penetration and its effects on fission product release.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Evaluation of Radiography for TRISO Buffer Layer Density Measurement

Tristructural isotropic (TRISO) fuel particles consist of a central uranium-bearing kernel and a series of coating layers designed to retain fission products and to ensure fuel performance. Several parameters such as thickness and density must be measured for these coating layers to show that they conform with fuel specifications. Current methods for measuring the density of pyrolytic carbon and silicon carbide layers (liquid gradient density column) and the buffer layer (mercury porosimetry) generate Resource Conservation and Recovery Act (RCRA) radiological-mixed waste. In addition, measurement of buffer and inner pyrolytic carbon layer densities require hot sampling or interrupted coating runs and the mercury porosimetry method used for buffer density measurement only measures the mean buffer density, not the interparticle distribution. A new approach has been evaluated to measure the density of coating layers in TRISO particles based on the dependence of x-ray attenuation in radiographs on material density. This method does not generate RCRA mixed waste, measures density on a particle-by-particle basis, and in principle is capable of measuring the density of all coating layers in a single process. Initial results using thinned TRISO particle sections to evaluate radiography measurement of density as a quality control characterization method are reported herein. In this work, the primary focus is on measurement of the density of the buffer layer; however, with appropriate calibration the method should be applicable to other coating layers. Improvements to the initial method and a full demonstration of the method on the remaining coating layers may be pursued as a future effort.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗